2. Methodology
2.1. Review Design and Scope
This review follows a critical narrative design supported by a transparent and structured literature search and study-selection process. The review brings together findings across different structural systems, data types, AI applications, validation approaches, and engineering outcomes to examine how AI-based methods contribute to structural condition assessment and rehabilitation. The literature search and study-selection process provides a clear and traceable basis for this synthesis while allowing evidence from a diverse body of research to be considered within the D2D framework.
The primary publication window was 2019–2026, corresponding to the rapid expansion of deep learning, connected sensing, and digital-twin research in structural engineering. Earlier studies were retained where they established an influential method or conceptual foundation for later work. The primary scope was existing building structures, including reinforced-concrete, steel, masonry, timber, and heritage systems. Infrastructure studies were retained selectively, where they supplied transferable evidence on SHM, prognosis, risk, or life-cycle maintenance decisions that remains comparatively sparse in the building-specific literature.
2.2. Literature Search and Study Selection
The literature search was conducted using Google Scholar, Scopus, and Web of Science and was last updated in July 2026. Searches combined structural terms with AI and decision terms. Structural terms included building structure, structural condition assessment, structural health monitoring, damage, deterioration, existing structure, and retrofit. AI terms included artificial intelligence, machine learning, deep learning, computer vision, neural network, ensemble learning, explainable AI, and physics-informed learning. Downstream terms included remaining useful life, risk, reliability, predictive maintenance, repair, strengthening, retrofitting, rehabilitation, and decision support.
One representative search string was: (“building structure” OR “structural condition assessment” OR “structural health monitoring” OR damage OR deterioration OR “existing structure” OR retrofit) AND (“artificial intelligence” OR “machine learning” OR “deep learning” OR “computer vision” OR “neural network” OR “ensemble learning” OR “explainable AI” OR “physics-informed learning”) AND (“remaining useful life” OR risk OR reliability OR “predictive maintenance” OR repair OR strengthening OR retrofit OR rehabilitation OR “decision support”).
The search was constructed in three layers. A core search identified review papers capable of mapping the principal terminology, data sources, and algorithm families. A second search targeted primary studies that advanced an output from one stage of the Data-to-Decision (D2D) Continuum to the next, for example, from component images to a building-level condition grade, or from a predicted damage state to residual safety. A third search targeted downstream terms that are frequently absent from AI-centered reviews, including intervention timing, maintenance policy, retrofit prioritization, life-cycle risk, and post-intervention monitoring. Citation chaining was then used to trace influential earlier studies that predated the main search window but remained necessary for interpreting current methods.
The corpus was not treated as a flat collection. Review articles were used to map the breadth and terminology of research streams, while primary studies were used to test the validity of specific engineering claims: a review stating that deep learning improves crack detection was not treated as equivalent to a field study that tested cross-building transfer, and neither was accepted as evidence that a repair decision had actually been improved. Distinguishing these evidence types limited the risk of circular citation, in which broad claims circulate by being repeatedly cited only by other reviews.
Peer-reviewed research and review articles were prioritized. A study was retained when it connected an AI method to at least one D2D stage and reported sufficient information to identify the structural context, input data, intended output, and evaluation approach. Pure architectural image classification, nonstructural energy management, and AI-based design optimization without a condition-assessment connection were excluded, as were duplicate records and documents lacking a traceable scholarly source.
Because this is a critical narrative review, retrospective systematic-review record counts and a formal study-flow diagram are not reported. The purpose of the structured search was conceptual coverage and traceable critical synthesis rather than exhaustive identification of every eligible record. The final corpus contained 82 retained sources, including 33 application-oriented sources used to examine the different stages and transitions of the D2D Continuum. The remaining sources included reviews, framework papers, and software or documentation sources used for field mapping and reproducibility context. This distinction prevents the corpus size from being interpreted as a systematic-review yield.
The first author conducted the initial screening and classification. Studies for which the scope, asset class, D2D assignment, validation setting, or evidence-maturity level was unclear were reviewed by both authors and resolved by consensus. No numerical quality score was used. Instead, relevance was assessed based on whether the structural context, input data, engineering output, and validation setting could be clearly identified.
A supplementary evidence matrix was prepared for the 33 application-oriented sources. The matrix records the structural system, data source, AI family, D2D stage, validation type, study-level evidence maturity, uncertainty treatment, decision relevance, principal limitation, and availability of public code or a ready-to-use tool. This provides a direct link between the individual studies and the synthesis presented in the review.
2.3. Evidence Extraction and Classification
Each retained source was classified by structural system, data modality, D2D stage, learning approach, validation setting, and reported engineering output. Validation settings were distinguished as simulated, laboratory, benchmark, field, or operational, since strong performance on balanced benchmark data does not by itself demonstrate robustness under rare damage, environmental variability, incomplete sensing, or cross-structure transfer.
Extraction also recorded the unit of analysis. Image patches, defects, members, stories, whole buildings, portfolios, and infrastructure networks represent different decision scales: a model evaluated on cropped crack images may be relevant to defect recognition, but it offers no aggregation rule for a building-level condition state, just as a regional post-disaster classifier may support emergency prioritization without resolving the repair needs of an individual frame. Recording scale explicitly kept performance reported at one level from being generalized, often implicitly, to another.
For quantitative results, each metric was recorded alongside its validation context rather than compiled into a single comparative table of accuracy values. Accuracy, precision, recall, intersection-over-union, mean average precision, error measures, and probabilistic calibration describe different properties and depend on class definition, prevalence, label quality, tolerance, and data partitioning. Direct numerical comparison was therefore made only when tasks and evaluation settings were sufficiently aligned; otherwise, the synthesis focused on what a study demonstrated and what remained untested.
Evidence was further classified using the interpretive maturity scale described in
Section 6.5: conceptual or simulated feasibility (E1), laboratory or curated-benchmark validation (E2), field data with study-specific validation (E3), prospective or multi-site operational validation (E4), and integration into a governed engineering workflow (E5). Each study was assigned the highest level for which the minimum expectations described in
Section 6.5 were supported by the reported evidence. If a study did not fully meet the requirements of a higher level, it was assigned to the lower level whose criteria were satisfied. These review-specific labels are qualitative, ordinal evidence categories rather than numerical scores or formal technology-readiness ratings.
2.4. Critical Synthesis Strategy and Limitations
The synthesis followed the stages of the D2D Continuum rather than structural type or algorithm family. For each stage, it examined the required input, the engineering meaning of the output, dominant AI approaches, validation practices, uncertainty, and connection to subsequent decisions; algorithm families were compared separately, and only after their stage-specific applications had been established.
Several limitations should be acknowledged. As a critical narrative review, the synthesis may be affected by database coverage, English-language and publication biases, the selected 2019–2026 emphasis, rapidly changing terminology, and disciplinary boundaries between SHM, earthquake engineering, reliability, asset management, and rehabilitation. The building-specific evidence base is smaller than the broader civil-infrastructure literature, particularly for prognosis and maintenance optimization. Negative results, deployment failures, inaccessible code, and proprietary datasets are likely underreported. Classification across D2D stages and E1–E5 categories also involves author judgment despite the stated criteria and consensus procedure. Rapid developments after the search period may alter the maturity of foundation models and autonomous agents; these technologies are therefore treated as emerging rather than established.
Publication volume was not treated as a proxy for readiness. A research area may appear mature simply because it has generated many publications; however, publication volume alone does not demonstrate engineering readiness. A densely published task may remain weakly validated if studies recycle the same benchmark or rely on simulation, whereas a small number of carefully designed field studies can provide stronger evidence for a bounded use case. The D2D maturity assessment therefore reflects depth of validation, connection to engineering outputs, and operational accountability, rather than the sheer volume of published work.
3. The Data-to-Decision (D2D) Continuum
3.1. Rationale and Boundaries
The D2D Continuum defines the chain of transformations required to convert observations into a justified intervention. Its starting point is not raw data alone, but data accompanied by provenance, calibration status, environmental context, and quality information. Its endpoint is not an automated repair command, but a decision package that combines predicted structural state, uncertainty, risk, feasible intervention alternatives, costs, constraints, and human approval.
The Continuum is deliberately task-centered rather than algorithm-centered. Algorithm taxonomies change quickly and often encourage a review to treat new architecture as contributions in themselves. Engineering tasks are more stable. A building still must be observed, damage must be distinguished from benign variability, structural significance must be established, future behavior must be considered, and action must be justified. Organizing the evidence by these tasks makes it possible to compare methods developed in different AI generations without losing sight of the decision they intended to support.
Table 1 positions this review against the principal organizing logic used in prior syntheses. Algorithm-oriented reviews compare model families, vision-oriented reviews organize evidence around image tasks, and SHM-oriented reviews commonly conclude at damage identification or condition monitoring. These perspectives remain valuable, but none makes the complete chain from traceable observation to authorized rehabilitation action its primary unit of analysis. Accordingly, the novelty of the proposed D2D Continuum lies not in introducing new AI algorithms, but in reorganizing the existing evidence according to the sequence of engineering decisions required to progress from structural observations to justified rehabilitation actions. Specifically, it is an interface-centered continuum that (i) specifies the engineering output required at each stage, (ii) exposes where uncertainty and responsibility pass between stages, and (iii) treats rehabilitation choice and post-intervention learning as integral parts of structural assessment rather than downstream afterthoughts.
The closest published frameworks overlap with different parts of the D2D Continuum. Decision-theoretic SHM frameworks use monitoring information, Bayesian updating, value of information, and utility to support inspection or maintenance decisions [
36,
37,
38]. Structural digital twins integrate sensing, simulation, learning, and management, including dynamic Bayesian updating and maintenance policies [
39,
40]. Life-cycle and bridge-management frameworks connect condition assessment with reliability, risk, and intervention planning [
37,
41,
42]. The D2D Continuum brings these related functions into an eight-stage, building-centered structure that follows the progression from structural observation to rehabilitation decision support. Its interface-centered organization defines the engineering output expected at each stage and makes uncertainty, validation requirements, and responsibility explicit as information moves between stages. Rehabilitation decisions and post-intervention learning are therefore treated as part of the same assessment continuum rather than as separate downstream activities.
The eight stages of the Continuum should not be interpreted as a requirement that one monolithic platform perform every operation. In practice, separate tools, organizations, and professional roles may own different stages. An inspection contractor may collect images, an SHM team may maintain sensors, a structural engineer may evaluate capacity, and an owner may choose among interventions. The Continuum instead identifies the interfaces at which information, uncertainty, responsibility, and assumptions must pass from one stage to another.
Figure 1 presents the Continuum. The central stages are linked by two cross-cutting layers. The first is a governance and assurance layer comprising data governance, uncertainty quantification, validation, explainability, and cybersecurity. The second is an engineering layer comprising mechanics, codes, reliability, constructability, and professional judgment. These layers prevent the pipeline from being interpreted as an autonomous sequence in which a single model directly determines a rehabilitation action.
3.2. Stages, Interfaces, and Outputs
Data acquisition produces quality-controlled observations from cameras, UAVs, accelerometers, strain sensors, acoustic emission, fiber optics, LiDAR, thermography, inspection records, and analytical models. Detection determines whether an abnormality or damage indicator is present. Localization determines where it occurs, while quantification estimates geometry, intensity, or severity. Condition and performance assessment translate these indicators into condition states, model parameters, capacity, serviceability, or demand-to-capacity measures.
Table 2 defines the engineering question, expected output, and principal validation concern for every stage.
Each interface requires an explicit output contract, the minimum information a stage must supply for the next stage, or a downstream engineer, to use it responsibly. Detection should provide more than a label; it should state the target class, probability or confidence, spatial and temporal coverage, and known failure modes. Localization should identify the coordinate system and structural component to which a result refers. Quantification should distinguish measured geometry from learned severity. Condition assessment should state whether the output is a rating, parameter estimate, limit-state measure, or professional conclusion. Without these contracts, downstream users may attach more engineering meaning to AI output than the model was trained to provide.
Prognosis estimates deterioration trajectories, threshold-crossing times, or remaining useful life. Risk assessment combines the probability of unfavorable states with their consequences. Rehabilitation decision support then compares feasible actions, including no action, further inspection, monitoring, repair, strengthening, retrofitting, use restriction, or replacement. The selected action changes the structure and should therefore be followed by post-intervention verification and model updating.
3.3. Uncertainty and Human Decision Gates
Uncertainty is transformed rather than eliminated as information moves through the pipeline. Measurement noise affects detection; domain shift affects localization and severity; model-form error affects capacity and prognosis; consequence assumptions affect risk; and uncertain costs or intervention effectiveness affect rehabilitation ranking. Passing a point estimate from one stage to the next conceals this accumulation. A decision-oriented implementation should propagate distributions, intervals, confidence measures, or scenario bounds wherever feasible.
Human decision gates are required after data-quality review, before a diagnosis is accepted as structurally significant, before safety or risk conclusions are issued, and before an intervention is authorized. Explainability should be matched to the gate. Pixel attribution may help verify a vision model, but a rehabilitation decision requires a higher-level explanation connecting evidence to structural mechanism, predicted consequence, and the trade-offs among alternatives [
43].
Four kinds of uncertainty are particularly important. Aleatory uncertainty reflects irreducible variability in loading, material properties, exposure, and deterioration. Epistemic uncertainty reflects incomplete knowledge, including limited training data and uncertain model form. Measurement uncertainty arises from sensors, calibration, imaging geometry, and labeling. Decision uncertainty concerns future costs, intervention effectiveness, stakeholder preferences, and consequences. A credible D2D implementation should identify which types are represented, which are neglected, and how sensitive the final action is to those choices.
The human role also changes across the pipeline. At acquisition, experts define coverage and verify data quality. At diagnosis, they test whether identified patterns are physically plausible. At performance assessment, licensed engineers reconcile AI outputs with analysis and code requirements. At the intervention stage, engineers and owners evaluate feasibility, constructability, occupancy, and risk tolerance. Human oversight is therefore not a generic statement placed after model development; it is a series of task-specific controls.
Uncertainty should be propagated through conditional models rather than appended only at the final stage. For example, a distribution for defect geometry should inform parameter updating; parameter and model-form uncertainty should then be carried into capacity or response estimates; deterioration and loading uncertainty should update reliability and remaining-life distributions; and these distributions should enter expected-risk or utility comparisons among interventions. Correlation between stage errors must be retained where the same measurements or simulations inform multiple stages. Sensitivity analysis should identify which upstream uncertainty can change the preferred action.
Before an AI output influences an intervention, the minimum evidence should include representative independent validation; documented provenance and leakage control; calibrated uncertainty or conservative bounds; a physically and code-consistent link to the claimed engineering output; predefined abstention and escalation rules; sensitivity of the action to model error; an auditable model/data version; and approval by the responsible engineer. High-consequence or irreversible actions additionally require external or prospective validation and an independent analytical or inspection check.
The proposed D2D Continuum provides the conceptual backbone for the critical synthesis presented throughout the remainder of this review.
4. AI Applications Across the Data-to-Decision (D2D) Continuum
This section follows the engineering workflow and is intentionally limited to the evidence, outputs, and stage-specific failure modes associated with each D2D transformation. Cross-cutting questions of transfer, validation, uncertainty, explainability, governance, and deployment are synthesized in
Section 6 rather than repeated at every stage.
4.1. Data Acquisition, Monitoring, and Fusion
Vision-based inspection is the most accessible route to scalable data collection because cameras can be deployed manually, on vehicles, or on UAVs. UAV research has demonstrated automated mapping in environments where direct access is difficult [
44,
45]. Engineering limitations are equally important: illumination, occlusion, camera distance, surface texture, and viewpoint can shift the data distribution, while visible damage may not represent hidden deterioration or residual capacity.
Acquisition design should begin with the engineering phenomenon, not the available device. Surface cracking and spalling may be captured with calibrated RGB imagery; moisture and delamination may require thermography or other nondestructive methods; global stiffness changes may require dynamic response; and corrosion or connection deterioration may require a combination of local and global measurements. A camera-only workflow is attractive because it is inexpensive, but choosing it by default can create a blind spot for internal or nonvisible damage.
Geometric traceability is essential for repeated inspection. Images should be linked to location, orientation, date, component, and acquisition settings. Photogrammetry, LiDAR, and BIM can provide a geometric reference, but registration errors may be comparable to the size of the defect being measured. If the objective is deterioration tracking, the system must distinguish a real change from differences in camera pose, focus, lighting, or surface condition.
Sensor-based SHM captures dynamic and quasi-static response through acceleration, strain, displacement, acoustic emission, temperature, and other variables. Fiber-optic sensing provides distributed or quasi-distributed measurements and is attractive for long-term monitoring, although installation, calibration, and interpretation remain structure-specific [
46]. IoT architectures support synchronized acquisition and remote access, but connectivity does not ensure data fitness. Missing values, timing errors, sensor drift, and abnormal data must be detected before downstream learning [
31].
Sensor placement is itself an inference problem. Dense arrays improve observability but increase cost, maintenance, bandwidth, and fault probability, while sparse arrays are practical but may leave several damage scenarios indistinguishable. AI studies often treat sensor layout as fixed and ideal, although deployment decisions determine what information reaches the model. Joint optimization of sensing and inference is therefore a promising direction, particularly where building access and occupancy constrain installation.
Long-term monitoring also requires separating damage from environmental and operational variability, since temperature, humidity, occupancy, equipment operation, wind, and boundary-condition changes can all shift response features. Baseline models trained over a narrow period may mistake seasonal change for damage. Robust systems therefore require sufficiently long baseline data, environmental normalization, change-point analysis, or models that explicitly represent these covariates.
Data fusion can combine complementary evidence and reduce dependence on a single sensor. Fusion may occur at raw-data, feature, model, or decision level. The choice is not purely computational: it determines how inconsistent evidence is reconciled and whether the contribution of each modality remains auditable. Current reviews identify fusion as a major opportunity but also note the scarcity of shared multimodal field datasets [
32].
Raw-level fusion can preserve detailed correlations but demands synchronization and compatible sampling. Feature-level fusion is more flexible but can hide how modality-specific preprocessing affects the result. Decision-level fusion allows independently validated models to be combined and may be easier to audit, although it can discard useful cross-modal relationships. Probabilistic graphical models are particularly attractive when evidence arrives asynchronously or has different reliability. The preferred strategy should be justified by the failure modes of the sensing system and the decision to be made.
Viewed as a body of evidence, acquisition studies demonstrate that increased sensing density does not by itself create decision-ready evidence. Their common strength is broader and more frequent observation; their common weakness is incomplete traceability from measurement quality to downstream error. Within D2D, acquisition is therefore an assurance stage rather than a neutral data source. Its output must be a quality-controlled observation package, because detection cannot distinguish damage from defects introduced by coverage, calibration, synchronization, or environmental variability.
4.2. Damage Detection
Damage detection has the largest and most mature AI evidence base. Image models identify cracks, spalling, corrosion-related manifestations, delamination signatures, and other surface defects. Vibration and strain models identify departures from a baseline state. Classical machine learning remains competitive when informative features and limited data are available, whereas deep learning is advantageous for high-dimensional images and complex signals.
The maturity of detection is uneven across defect types. Crack datasets are relatively abundant, and labeling is visually intuitive, which has encouraged rapid model development. Corrosion, delamination, connection damage, and hidden deterioration are less represented and often require indirect sensing. Even for cracks, differences in material texture, paint, repair history, contamination, and image scale can cause models to confuse nonstructural patterns with damage.
Unsupervised and one-class approaches address the practical reality that healthy-state data are more available than labeled damage. They learn baseline representation and flag departures, but an anomaly is not necessarily damaged. Changes in occupancy, sensor configuration, or environmental exposure can produce the same statistical signal. These methods are most defensible as screening tools that trigger review or additional inspection, not as standalone diagnoses.
Autoencoder-based damage identification demonstrated early that representation learning could extract damage-sensitive structure from noisy response data [
47]. More recent work uses deep metric learning to place similar damage conditions close in a learned space, enabling localization without a rigid set of predefined classes [
48]. These advances improve flexibility, but their simulated-to-real transfer remains governed by the fidelity of the numerical model and the similarity of operational conditions.
Reported accuracy can be misleading when the negative class is uniform, damage examples are oversampled, or images from the same structure appear in both training and testing sets. For rare safety-relevant damage, sensitivity, false-negative rate, calibration, and precision-recall behavior are more informative than overall accuracy. Cross-site validation is particularly important because background texture or acquisition conditions can become unintended shortcuts.
Dataset partitioning deserves special attention. Random image-level splitting can place adjacent crops, frames from the same video, or repeated views of one defect in both training and testing subsets. The model then demonstrates recognition of a scene rather than transferring to a new building. Structure-level and event-level holdouts are more demanding and more informative. Prospective testing, in which data is collected after model development, provides an even stronger test of operational validity.
The central lesson from detection research is that AI can screen large image and signal collections with useful sensitivity under bounded conditions. It does not establish that a detected pattern is structurally consequential, and studies disagree mainly because targets, negative cases, and validation units are not standardized. The D2D implication is precise: detection supplies an evidence flag, not a diagnosis. The next stage is required because a flag without a defensible physical location cannot be reconciled with a component, load path, or inspection action.
4.3. Damage Localization
Localization ranges from identifying a structural region to generating pixel-level damage maps. Object detectors provide bounding regions, while semantic and instance segmentation estimate spatial extent. For vibration-based SHM, localization may be inferred from spatial patterns of modal changes, transmissibility, or sensor responses. The localization scale must match the intended action: identifying a damaged story may support emergency screening, whereas repair planning may require member- or defect-level resolution.
Localization outputs should be evaluated geometrically as well as statistically. Intersection over-union is useful for segmentation, but repair planning may depend on maximum crack extent, proximity to a connection, or continuity across a critical region. A model can achieve a favorable average segmentation score while missing a small but structurally important branch. Evaluation should therefore include defect-level recall and tolerance measures tied to the intended inspection scale.
For vibration-based methods, localization is an inverse problem: different combinations of damage, loading, and boundary change can produce similar response patterns. Physics-guided learning can restrict the solution space. Residual neural networks constrained by structural dynamics have shown improved localization and quantification under limited training data [
49]. Physics-encoded self-supervised learning similarly seeks to recover location and severity without requiring labeled damaged-state measurements [
50]. These methods are promising because they respond directly to scarcity, although their experimental evidence still represents a narrower range of structures than field deployment requires.
UAV mapping and segmentation have strengthened the geometric connection between detections and physical locations. Nevertheless, image coordinates must be registered to structural components, and repeated observations must be distinguishable from duplicate views of the same defect. BIM or digital-twin integration can help maintain this identity, but automated registration in field conditions remains a practical bottleneck.
From an engineering perspective, localization marks the shift from recognizing damage to assigning it a structural identity. Agreement is strongest for pixel- or region-level mapping in controlled imagery; evidence is weaker for registration across repeated field inspections and for inverse localization from sparse response data. In D2D terms, localization is complete only when the result can be referenced to a persistent component and coordinate system. That requirement creates the basis for quantification, where spatial evidence must be converted into a measurable extent or severity.
4.4. Damage Quantification and Severity
Quantification estimates crack width and length, damaged area, corrosion level, stiffness loss, or a severity class. It is a more demanding task than detection because a visually correct segmentation does not guarantee metrically accurate geometry. Scale calibration, perspective correction, surface topology, and uncertainty in defect boundaries directly affect the estimate.
The distinction between measurement and inference should remain visible. Crack width derived from a calibrated image is a measurement affected by resolution and geometry. Stiffness loss inferred from dynamic response is a model-dependent parameter. A damage grade predicted from an image is a learned category whose meaning depends on the labeling protocol. Reporting all three as “severity” conceals important differences in traceability and validation.
Severity labels may be convenient for model training but can mix appearance and structural consequences. A wide surface crack is not necessarily more critical than a narrow crack at a vulnerable detail. Quantification should therefore state whether the target is physical geometry, a condition category, or a proxy for performance. Models that infer stiffness reduction or post-event safety categories move closer to structural meaning, as demonstrated in post-earthquake reinforced-concrete assessment research [
51].
Probabilistic quantification is preferable where labels or measurements are uncertain. Recent image-based work on reinforced-concrete beams has combined crack features, probabilistic modeling, and explainable AI to relate visible patterns to displacement- and strength-oriented damage indicators [
52]. The value of such an approach lies less in attaching an explanation graphic to a classifier than in expressing how uncertain visual evidence maps to engineering-relevant quantities.
Here, the central disagreement is not whether AI can estimate a number, but what that number represents. Geometric measurements inferred stiffness changes, and expert-defined severity classes have different evidentiary status and should not be compared as interchangeable outputs. D2D resolves this ambiguity by requiring quantification to declare its unit, calibration basis, and uncertainty. Only then can the result enter condition assessment without being assigned more structural meaning than the evidence supports.
4.5. Structural Condition and Performance Assessment
Condition assessment integrates multiple observations into a statement about present structural state. AI applications include condition ratings, anomaly classification, parameter identification, surrogate modeling, model updating, and rapid safety screening. The crucial distinction is between recognizing a damage pattern and demonstrating its implication for strength, stiffness, stability, serviceability, or robustness.
Three routes are apparent in the literature. The first aggregates component-level observations into a building-level score or grade. The second maps structural response or damage patterns to performance measures such as residual capacity or a safety state. The third updates parameters of an analytical model and uses that model to assess performance. The routes should not be conflated. Weighted aggregation may be appropriate for rapid screening, whereas a safety-critical occupancy decision requires a demonstrated relationship to structural behavior.
Post-earthquake research illustrates both progress and remaining difficulty. A machine-learning framework has mapped response and component-damage patterns to residual collapse capacity and a probabilistic safe/unsafe state for a reinforced-concrete building [
53]. Other work has linked reconstructed hysteretic behavior and damage indices to experimentally observed damage states [
54]. These studies are more decision-relevant than image classification alone because their targets have explicit performance meaning.
Image-based systems are also moving beyond isolated components. A component-recognition workflow has aggregated detected damage into an overall post-earthquake grade for concrete-building portfolios using code-informed weighting [
55]. Transfer-learning systems have been tested for preliminary damaged-building assessment in Taiwan [
56], while ensemble segmentation has used satellite imagery and field-derived labels from the 2023 Türkiye earthquakes for large-scale rapid assessment [
57]. These tools support emergency prioritization, but their spatial scale and visible-damage basis mean that they supplement rather than replace detailed structural evaluation.
Hybrid rule-based and deep-learning approaches are particularly relevant to D2D because they make aggregation assumptions explicit. A recent framework for reinforced-concrete buildings combines object detection with engineering rules and meta-level fusion to improve post-earthquake assessment across image datasets [
58]. The approach demonstrates a useful design principle: when the decision logic is partly codified, it may be safer to expose those rules than to require a black-box model to learn them implicitly.
Hybrid approaches are promising because they constrain learning with structural mechanics or use AI as a surrogate within an analytical assessment. Purely data-driven models can interpolate efficiently within a represented domain but may produce physically inconsistent results when loading, geometry, or boundary conditions change. Model updating and digital twins provide a mechanism for reconciling observations with a physical model, although uncertainty in both the measurements and the model must be acknowledged.
Surrogate modeling can reduce the computational burden of nonlinear analysis and support portfolio-scale screening. Machine-learning models have been used to predict seismic response and performance levels of reinforced-concrete buildings from structural and ground-motion descriptors [
59]. Deep-learning models have also considered cumulative mainshock-aftershock damage and pre-existing conditions [
60]. The benefit is speed; the risk is that a surrogate may be applied outside the geometry, detailing, material, or hazard range represented during training.
Real-time assessment introduces another temporal dimension. Multi-source frameworks can combine building characteristics, ground-motion intensity measures, and monitoring records to estimate how a damage state evolves during an event [
61]. Such systems may support emergency decisions, but real-time speed does not remove the need for conservative thresholds, out-of-domain detection, and confirmation after the event.
Condition-assessment evidence provides the most important conceptual bridge in the Continuum: they attempt to translate observed damage into performance meaning. The strongest studies define targets such as residual capacity, limit-state probability, or a code-relevant safety category; weaker studies relabel visual severity as structural condition without demonstrating that connection. For practice, D2D therefore requires an explicit analytical or empirical mapping from quantified evidence to capacity, serviceability, or safety before prognosis is attempted.
4.6. Prognosis and Remaining Useful Life
Prognosis shifts the question from present state to future evolution. Relevant outputs include the deterioration rate, the probability of reaching a limit state within a time horizon and remaining useful life (RUL). Methods include regression, survival models, recurrent networks, Bayesian updating, degradation models, and hybrid physics-data approaches. Recent reviews report growing interest in AI-based RUL for civil infrastructure but emphasize limited failure histories, censoring, changing exposure, and weak transfer across assets [
62].
Civil structures create a different prognostic problem from replaceable mechanical components. Buildings generally do not have a large population of identical components operating under controlled duty cycles. Their deterioration is influenced by construction variability, repairs, modifications, exposure, occupancy, and low-frequency extreme events. Failure data are sparse partly because inspection and intervention prevent many assets from reaching an observed end state. Models trained on condition ratings may therefore learn inspection and maintenance practice as much as physical deterioration.
Prognostic targets must be defined carefully. Remaining time to a condition-rating threshold is not the same as remaining time to a structural limit state. A rating may be ordinal and partly subjective, whereas a limit state is tied to performance. For decision support, it may be more useful to predict the probability of crossing several thresholds under alternative inspection and intervention scenarios than to report a single RUL value.
RUL should not be interpreted as a single deterministic date. For buildings, degradation may be localized, episodic, or dominated by uncertain future hazards and use. A more defensible output is a probability distribution or scenario-conditioned interval tied to a defined performance threshold. Prognostic validity also requires temporal validation: random splitting of observations from the same degradation history can leak future information into training.
Maintenance actions complicate prognosis because they alter the process being predicted. A model trained in untreated deterioration may overstate future damage after repair, while a model trained on historical records may embed inconsistent intervention quality. Prognosis and maintenance should therefore be coupled: the state-transition model should represent how different actions change conditions and uncertainty. This coupling is essential if RUL is to support rehabilitation rather than remain a descriptive forecast.
Prognostic models demonstrate useful predictive capability but remain methodologically fragmented by inconsistent end states, censoring treatment, and intervention histories. Reported RUL values are therefore less comparable than their common terminology suggests. Within D2D, prognosis should produce a conditional distribution tied to a stated threshold and action history, not a single asset-life estimate. This probabilistic output is what enables the subsequent risk stage to combine likelihood with consequence.
4.7. Reliability and Risk Assessment
Reliability and risk provide a bridge between predicted condition and intervention urgency. Reliability concerns the probability of violating a limit state, whereas risk additionally considers consequences to occupants, functionality, economic value, heritage, and recovery. AI can accelerate surrogate evaluations, recognize post-event damage states, and update risk as evidence arrives. It should not obscure the assumptions embedded in hazards, limit states, consequence models, or acceptable risk.
Machine learning is useful in structural reliability where repeated nonlinear analyses make direct probability estimation expensive. Surrogates can approximate demand, capacity, or limit-state functions, but approximation error may be amplified in the low-probability tail that controls safety decisions. Research on machine-learning approximations of life-cycle reliability and risk shows that error analysis must be performed at the probability and risk level, not only at the response-prediction level [
63].
Risk assessment also provides a principled basis for prioritization. Two defects with similar predicted severity may warrant different actions because occupancy, redundancy, failure mode, and consequence differ. Conversely, a high probability of a minor serviceability issue may not dominate a lower-probability brittle failure. AI models should therefore avoid collapsing risk into an opaque score unless the utility, consequence, and tolerance assumptions are available for review.
Risk-informed use also changes the evaluation of model error. A false negative near a critical limit state is not equivalent to a false positive in a low-consequence component. Cost-sensitive learning and decision-theoretic metrics are therefore more relevant than unweighted classification accuracy. Uncertainty calibration is essential because confidence values may be used to trigger inspections, restrictions, or emergency action.
Decision thresholds should be evaluated through expected consequences and professional requirements. A screening model may be tuned for high sensitivity and accept more false alarms; an occupancy recommendation may require a conservative abstention region; and an intervention-ranking model may need stability under alternative cost and consequence assumptions. Reporting one operating threshold without sensitivity analysis limits the usefulness of the result.
Risk-oriented evidence makes clear why predictive accuracy alone is an insufficient endpoint. Models that perform similarly on average can imply different actions once failure consequences, risk tolerance, and asymmetric errors are considered. The unresolved gap is the consistent propagation of calibrated uncertainty from diagnosis and prognosis into reliability and consequence models. In the D2D chain, risk is the decision interface that converts uncertain future conditions into urgency and provides the defensible basis for comparing intervention alternatives.
4.8. Rehabilitation, Retrofitting, and Maintenance Decision Support
This is the least developed stage of the current evidence base. Most AI studies support inspection or diagnosis; relatively few evaluate the selection of structural repair or retrofit strategies. Building research has demonstrated decision-support models that predict damage causes and recommend maintenance solutions, showing that downstream integration is feasible [
64]. Broader structural-maintenance research has also used reinforcement learning to optimize actions under life-cycle cost and safety objectives [
65].
The building-maintenance literature provides additional evidence on resource allocation. Association-rule mining has been used to identify accelerated deterioration relationships among building components and support maintenance, rehabilitation, and repair planning [
66]. Such approaches address a practical feature often absent from component-level AI studies: interventions compete for limited budgets and may affect interdependent building systems.
Reinforcement learning formalizes maintenance as a sequence of state observations, actions, transitions, and rewards. This is conceptually aligned with the closed-loop D2D Continuum because actions alter the future state and new observations update the policy. Recent decision-making agents have generated maintenance schedules and budget information across infrastructure components [
67]. However, simulation-based success depends on the transition model, action effectiveness, reward weights, and representation of rare safety events. A policy can be computationally optimal and still be unacceptable if those assumptions are incomplete.
A rehabilitation decision has at least four components: whether intervention is needed, when it should occur, where it should be applied, and which action should be selected. Feasible alternatives must respect structural mechanisms, code requirements, constructability, occupancy, cost, service disruption, intervention durability, embodied carbon, and heritage constraints. An AI recommendation that ignores feasibility may optimize an abstract objective while producing an unusable action.
The intervention set should be defined by engineering practice before optimization. Depending on the condition, alternatives may include no action, targeted investigation, monitoring, load restriction, local repair, corrosion mitigation, section enlargement, jacketing, fiber-reinforced polymer strengthening, connection modification, supplemental damping, seismic retrofit, partial replacement, or decommissioning. The applicability and effect of each option depend on material, load path, damage mechanism, access, compatibility, and code objectives. AI can assist the comparison and ranking of feasible alternatives, but it cannot infer an unconstrained intervention vocabulary safely from historical labels alone.
Multi-criteria decision support is more appropriate than a single predicted “best” action. Safety is a constraint and an objective; cost, carbon, disruption, heritage value, durability, and resilience introduce additional trade-offs. Recommendations should show how rankings change with stakeholder weights and uncertain intervention performance. A robust option that remains acceptable across scenarios may be preferable to an apparently optimal option whose ranking is highly sensitive.
The appropriate near-term role is decision support rather than autonomous authorization. AI can screen alternatives, estimate outcomes, update priorities, and identify cases requiring expert review. The engineer remains responsible for confirming the structural model, intervention details, load path, compatibility, and compliance. Post-intervention monitoring should test whether the assumed benefit was achieved and provide data for updating future recommendations.
The limited direct evidence for structural rehabilitation should not be obscured by borrowing examples from energy retrofit or generic facility management. Those fields provide useful methods for multi-criteria optimization and stakeholder interaction, but structural interventions are governed by load paths, limit states, detailing, and life-safety obligations. Transfer is appropriate at the decision-architecture level, not as proof of structural validity.
The evidence at this final stage is markedly less mature than the evidence for detection. Existing studies demonstrate that AI can rank options or optimize policies under defined models, but they rarely validate whether the recommended structural intervention was feasible, implemented, and effective. This imbalance is a central finding of the review rather than a simple research shortage. D2D makes it visible by requiring rehabilitation support to return a ranked, constraint-compliant action set with uncertainty, rationale, and human authorization.
4.9. Closing the Loop
The D2D process becomes a learning system only when outcomes return to the evidence base. After repair or retrofit, baseline measurements should be re-established, model parameters updated, and residual anomalies investigated. Intervention records should preserve what was done, where with, which materials, under what assumptions, and with what observed effect. Without this feedback, maintenance models learn from inspections but not from the success or failure of actions.
Digital twins provide one possible architecture for this loop. A probabilistic digital-twin framework can assimilate sensor data, update structural state, represent uncertainty, and inform maintenance actions through a repeated observations-to-decisions process [
39]. This is more substantive than a static 3D model labeled as a digital twin. The defining feature is an operational connection that updates state and supports a decision.
The distinction between a digital model, a digital shadow, and a decision-capable twin matters. Many proposed twins visualize data or update parameters but do not demonstrate predictive validity or action selection. Emerging AI-enabled frameworks combine sensor fusion, Bayesian inference, health indices, and maintenance support [
68]. Their promise is high, but field evidence, governance, and building-specific application remain limited.
Closing the loop changes the interpretation of every preceding stage. Without post-intervention observations, the Continuum can learn correlations between data and historical labels but cannot learn whether its decisions improved safety, service life, cost, or disruption. In D2D, feedback is therefore not an optional digital-twin feature; it is the mechanism that tests the intervention hypothesis and updates both the asset model and the evidence base.
The evidence across these stages establishes a boundary between workflow synthesis and implementation assurance.
Section 4 has identified what each stage must produce and where stage-specific errors arise. Whether those outputs transfer across structural systems remain valid outside the development dataset, adequately communicate uncertainty, and can be governed in practice are separate translational questions addressed in
Section 6.
5. Selecting AI Paradigms for Decision-Relevant Structural Tasks
5.1. Selection Criteria and Established Paradigms
Within the Data-to-Decision (D2D) Continuum, the appropriate AI paradigm is determined by the engineering output and its intended use, not by architectural novelty. Model selection should begin with five questions: What quantity or state must be inferred? What evidence is available at deployment? How far must the model extrapolate beyond its training domain? What uncertainty and explanation must accompany the output? What are the consequences of error? These questions distinguish a screening classifier from a capacity-assessment surrogate or a maintenance-policy model even when all three use related computational tools.
Data volume and dimensionality matter, but they are not sufficient selection criteria. A highly parameterized model can be justified for diverse image collections, whereas a lower-complexity model may be more defensible for a small tabular condition dataset. Where mechanics governs extrapolation, a hybrid model may be preferable despite higher implementation effort. Where decisions are sequential, a policy model may be appropriate only if its reward, constraints, and fallback behavior can be audited.
Support vector machines, random forests, gradient boosting, and shallow neural networks remain well suited to structured datasets and engineered features. Their lower data demand and computational cost can favor embedded or interpretable applications. Ensembles can improve robustness but do not automatically solve domain shift or uncertainty calibration. Their principal limitation is dependence on features that may not transfer across structures or sensing configurations.
Conventional machine learning is sometimes presented as a historical stage superseded by deep learning, but that characterization is misleading. For the modest, structured datasets common in engineering, tree ensembles and support vector machines remain efficient, inspectable baselines against which added model complexity should be justified.
The main risk is that engineered features encode site-specific assumptions. Modal features depend on excitation, sensor placement, environmental normalization, and identification quality. Condition ratings may reflect organizational practice. A model can be transparent in mathematical form while remaining opaque in its data-generating assumptions. Interpretability therefore requires documentation of features and acquisition, not merely selection of an algorithm considered interpretable.
Figure 2 organizes the principal AI families by the data types for which they are commonly appropriate. The mapping is indicative: images may be converted to tabular descriptors, signals may be represented as spectrograms, and hybrid models may combine several modalities. The diagram therefore links families to dominant data-generating structures rather than imposing exclusive categories.
Figure 3 summarizes the stage-dependent fit of the principal AI paradigms. The relative fits reflect a qualitative synthesis of the reviewed literature rather than a quantitative comparison of model performance. Data acquisition is addressed separately in
Section 4.1 and is therefore not repeated here, since paradigm choice at that stage is governed primarily by sensing hardware and coverage rather than by learning architecture. The purpose of
Figure 3 is comparative rather than prescriptive: the changing pattern across D2D stages reinforces that architecture choice should follow the required engineering output.
5.2. Common Model-Training and Evaluation Practices
Across the application literature, supervised models are commonly trained after cleaning, normalization, feature engineering or augmentation, followed by a training-validation-test partition. Hyperparameters are selected by grid, random, or Bayesian search, and early stopping or regularization is used for high-capacity models. Class weighting, focal loss, oversampling, and targeted augmentation are frequent responses to rare damage classes. Transfer learning is common for image models, whereas simulation-assisted learning and pretraining on undamaged monitoring archives are common when labeled field damage is scarce.
The dominant weakness is not the absence of training sophistication but the evaluation unit. Randomly splitting correlated image patches, overlapping signal windows, or multiple simulations from the same structural model can leak structure-specific information into the test set. More defensible practice holds out complete components, structures, sites, events, or time periods; performs all preprocessing and feature selection within the training folds; reports class prevalence and repeated-resampling variability; compares against simple baselines; and calibrates probabilities separately from optimizing discrimination.
Reproducibility requires reporting random seeds, data partitions, preprocessing, augmentation, architecture, loss, optimizer, learning-rate schedule, stopping rule, hyperparameter search space, selected configuration, hardware, software versions, and evaluation code. Common environments include Python with NumPy, pandas, SciPy, scikit-learn, TensorFlow/Keras or PyTorch; gradient-boosting studies frequently use XGBoost, LightGBM, or CatBoost; OpenCV supports image processing; and SHAP is widely used for model explanation [
69,
70,
71,
72,
73,
74,
75,
76]. These libraries facilitate development but do not by themselves establish engineering validity.
5.3. Deep and Temporal Models for High-Dimensional Evidence
Deep learning dominates image-based inspection because it learns hierarchical representations and supports detection, segmentation, and multimodal processing. Its disadvantages are high data demand, sensitivity to hidden shortcuts, computational burden, and limited transparency. Transfer learning reduces training requirements, but pretraining on natural images does not guarantee sensitivity to subtle structural defects.
Architecture choice should follow the output. Image classifiers are appropriate when the entire image has one label; object detectors identify multiple bounded defects; segmentation models estimate extent; and vision transformers may capture longer-range context. Using a classifier on cropped defects can produce impressive results while avoiding the harder task of finding damage in a full inspection scene. The reported task must therefore be examined before the architecture is credited with practical inspection capability.
Temporal deep learning includes recurrent networks, temporal convolution, attention, and transformer-based sequence models. These methods can represent nonlinear dependencies in vibration or strain histories, but their apparent predictive skill can be inflated by overlapping windows and temporal leakage. Evaluation should hold out complete time periods, events, or structures. For long-term SHM, the ability to adapt without forgetting previously learned normal states is also important.
5.4. Physics-Guided and Probabilistic Paradigms
Hybrid models combine physical structure with learned components. Examples include surrogate models constrained by mechanics, residual learning around an analytical prediction, simulation-assisted training, and digital-twin updating. These approaches are attractive where data are sparse and extrapolation matters. Their success depends on whether the physical assumptions are appropriate; an incorrect constraint can introduce systematic bias while giving an appearance of credibility.
Physics can enter learning through data generation, architecture, loss functions, or hybrid coupling. Simulation-generated data require the reality gap to be quantified; constrained architectures and losses encode governing relationships; and solver–learner hybrids estimate uncertain parameters or residuals. These strategies offer different balances of interpretability, flexibility, and computational cost.
Physics-guided residual networks demonstrate how governing equations can reduce data demand in damage identification [
49]. Self-supervised physics-encoded models attempt to recover damage without labeled damaged-state records [
50]. Their strongest contribution is not guaranteed extrapolation, but a more defensible inductive bias. They still require validation against modeling errors, unrepresented damage mechanisms, and field boundary conditions.
Probabilistic and Bayesian approaches address a different requirement: representing uncertainty explicitly rather than constraining the model with mechanics. Bayesian updating, Gaussian processes, and probabilistic graphical models estimate a distribution over the target rather than a single value, which makes them well suited to prognosis, reliability updating, and digital-twin state estimation, where a defensible interval or probability is the required engineering output rather than a convenience [
63,
68]. Their principal cost is computational: online updating and calibrated inference can be expensive at the scale of a building portfolio, and their interpretability is high for the uncertainty itself but model dependent for the underlying physical or statistical assumptions.
5.5. Data-Scarce and Emerging Paradigms
Self-supervised learning may exploit large unlabeled monitoring archives, and generative models may augment rare damage conditions. Foundation models offer reusable representations and natural-language interfaces, but evidence for safety-critical structural assessment is still limited. Synthetic data must be evaluated for physical and visual fidelity, and language-model outputs require retrieval, traceability, and professional verification. These approaches should be assessed through demonstrated engineering benefit rather than novelty.
Self-supervised objectives can learn representations through reconstruction, contrast, prediction of masked segments, or consistency across augmented views. The design of augmentation is critical. An augmentation that changes crack width, removes a high-frequency response, or alters temporal order may destroy the very feature that has structural meaning. Domain knowledge is needed to define transformations that represent nuisance variability without changing the label.
Generative models can create rare damage images or response histories, but visual realism is insufficient. Their value should be demonstrated on held-out real data, with evidence that synthetic samples preserve defect geometry, material interaction, sensor noise, and structural consequence and improve transfer rather than only benchmark performance.
Reinforcement learning occupies a related but distinct position among emerging paradigms: rather than estimating a state or distribution, it formalizes maintenance and inspection as a sequence of actions, transitions, and rewards, an approach already illustrated by RL-based maintenance-policy optimization and infrastructure decision-making agents in
Section 4.8 [
65,
67]. Its data demand is high because it typically requires a simulator or extensive longitudinal decision histories, and its policy rationale often needs additional explanation before it can be audited at a safety-critical decision gate.
Foundation models may support annotation, retrieval, report drafting, coding, and multimodal representation. Large language models may help engineers find records or explain established analyses, but unconstrained generation is incompatible with traceable safety decisions. Retrieval-augmented and tool-bounded systems should cite the governing data, model, and code provisions behind an answer. Agentic workflows remain an emerging research direction until their actions, permissions, failure recovery, and human approvals are demonstrably controlled.
5.6. Synthesis: No Universally Superior Paradigm
No model family is universally superior. The preferred approach depends on the decision target, data volume, dimensionality, required explanation, computational environment, and consequences of error. Classical models may be preferable for low-dimensional condition indicators; deep vision models for defect mapping; temporal models for prognosis; hybrid models for capacity assessment; probabilistic and Bayesian models where a calibrated interval is itself the required output; and reinforcement learning for sequential maintenance policies. The comparison in
Table 3 emphasizes suitability and evidence maturity rather than ranking by isolated accuracy.
Model selection should also account for the cost of maintaining the model. A marginal accuracy improvement may not justify a pipeline that requires specialized hardware, continuous annotation, or frequent retraining. In a building portfolio, simple models with clear abstention rules may be easier to govern than a single complex model expected to cover many materials and damage mechanisms. Modularity also allows individual components to evolve as evidence matures, without requiring reconstruction of the entire D2D Continuum.
6. Translation to Engineering Practice
Stage-level capability and appropriate model choice within the Data-to-Decision (D2D) Continuum are necessary but insufficient for professional use. This section therefore moves from what AI can infer to the conditions under which those inferences can be trusted, transferred, integrated, and governed in real assessment and rehabilitation workflows.
6.1. Data Quality, Representativeness, and Benchmarking
Structural damage data are scarce because serious damage is uncommon, access is restricted, and labeling requires expertise. Available datasets often overrepresent clean, visible cracks and underrepresent ambiguous or hidden damage. Benchmark datasets support reproducibility but can encourage optimization to a narrow distribution. Useful benchmarks should include acquisition metadata, environmental variation, negative examples, uncertainty in labels, and held-out structures.
Numerical simulation and field data serve complementary roles. Nonlinear time-history analysis (NLTHA), finite-element models, and digital twins can generate balanced labels, expose rare or unsafe damage states, isolate mechanisms, and support sensitivity studies at relatively low marginal cost. They also provide quantities that are difficult to observe directly, such as plastic rotations, residual capacity, or limit-state exceedance. Their limitations are the reality gap: uncertain material and damping models, idealized boundaries, incomplete nonstructural interaction, simplified deterioration, and synthetic noise may allow a learner to exploit artifacts that do not exist in practice.
Field observations capture real geometry, construction variability, occupancy, environmental effects, sensor failures, mixed damage mechanisms, and operational constraints, making them indispensable for external validation and deployment claims. They are nevertheless sparse, imbalanced, incompletely labeled, and sometimes confounded by inspection urgency and inaccessible pre-event baselines. At D2D acquisition through localization, field data should dominate validation, with simulation used for augmentation and controlled testing. At quantification and performance assessment, calibrated simulation can supply latent engineering labels but should be conditioned and checked using experiments or field observations. At prognosis, risk, and rehabilitation stages, simulation is convenient for scenario exploration and policy training, but transition models, costs, consequences, and recommended actions require field records, expert review, and prospective outcome monitoring before decision use.
Label quality is a structural-engineering issue. Crack masks can vary among annotators, damage grades can depend on inspection guidance, and performance states may require analysis rather than observation. Consensus labels hide disagreement unless uncertainty is retained. For ambiguous cases, soft labels, multiple annotations, or adjudication records provide more information than a single definitive class.
Dataset documentation should include the structural system, material, age, exposure, acquisition device, image scale, sensor layout, sampling, preprocessing, labeling protocol, class prevalence, and permitted uses. Without this information, users cannot determine whether a dataset represents their building stock. Data sheets and model cards can make these boundaries explicit.
6.2. Cross-System Interpretation for Existing Buildings
The D2D stages are common across structural systems, but the observable damage, useful data, and intervention logic are not. Reinforced-concrete, steel, masonry, timber, and heritage buildings should therefore not be pooled without considering material behavior and structural mechanism. A model that transfers visually across surfaces may still fail to transfer structurally because the same apparent feature has different implications.
Reinforced-concrete buildings dominate vision-based research. Cracks, spalling, exposed reinforcement, and corrosion staining provide visible targets, and post-earthquake reconnaissance supplies damage-state imagery. The main unresolved issue is the mapping from appearance to mechanism. Flexural, shear, shrinkage, thermal, settlement, and corrosion-related cracks may overlap in appearance while differing greatly in consequence. Component type, orientation, reinforcement detailing, demand, and crack evolution are needed to interpret the image. Post-earthquake studies that aggregate component damage or related response patterns to residual safety represent important progress because they incorporate some of this context [
53,
55].
For steel buildings, connection behavior, fatigue cracking, corrosion, local buckling, residual deformation, and fire-related changes are important targets. Surface images may reveal corrosion and gross deformation, but small fatigue cracks and connection distress often require targeted nondestructive evaluation. Vibration-based models can identify global changes, yet nonstructural components and occupancy may strongly affect building response. Physics-guided and model-updating approaches are therefore particularly relevant, provided the analytical model represents connection and boundary behavior with sufficient fidelity.
Masonry buildings pose a different problem because material heterogeneity, construction history, moisture, prior repair, and irregular load paths complicate both modeling and labeling. Cracking may follow mortar joints, units, interfaces, openings, or settlement patterns. Heritage masonry adds constraints associated with minimal intervention, reversibility, fabric conservation, and incomplete documentation. Reviews of damage detection in heritage masonry show growing use of imaging and monitoring but also emphasize the difficulty of obtaining generalizable datasets and ground truth [
77]. For these structures, the D2D decision package must incorporate heritage significance and intervention compatibility in addition to structural risk.
Timber buildings remain underrepresented in the reviewed AI corpus. Relevant deterioration includes moisture-related decay, biological attack, connection degradation, delamination in engineered wood products, and fire effects. Many critical conditions are not visible on the surface, and material variability complicates universal thresholds. This underrepresentation should be treated as a research gap, particularly as mass-timber construction expands and existing timber stocks require long-term assessment.
The structural system also affects the appropriate unit of analysis. A crack segment may be meaningful in reinforced concrete, a connection or brace in steel, a wall panel or pier in masonry, and a member–connection–moisture zone in timber. Dataset and model design should follow these units. A generic “damage image” class obscures the relationship between an observation and the element that carries load.
The evidence from bridges and other infrastructure remains useful when transfer is explicit. Bridge SHM offers experience with continuous monitoring, environmental normalization, sensor calibration, deterioration modeling, and maintenance optimization [
20,
33]. Buildings can benefit from these methods, but differences in loading, ownership, access, nonstructural interaction, and alteration history require revalidation. Transferable methodology is not transferable performance.
To make the evidentiary basis of these transfer considerations explicit, the 33 application-oriented sources were grouped according to their structural context: 17 concern buildings or building components, 9 concern bridges or other infrastructure assets, and 7 use generic structural specimens, synthetic systems, or cross-asset methods. Building evidence is concentrated in post-earthquake image assessment, member-capacity prediction, and component maintenance, while bridge and infrastructure studies contribute more strongly to long-term monitoring, reliability, prognosis, and maintenance optimization. Conclusions from the first group are treated as directly supported by building evidence, whereas conclusions drawn from the second and third groups are identified as transferable hypotheses that require building-specific revalidation.
Across all systems, the most reliable research strategy is to combine population-level learning with building-specific updating. Population data can provide priors, pretrained representations, and common failure patterns. Building-specific observations, drawings, tests, and models can then adapt those priors to the asset. This hierarchy is more realistic than expecting one universal model to cover every material, structural configuration, age, and environment.
6.3. Generalization, Domain Shift, and Robustness
Generalization across buildings is constrained by differences in material, age, detailing, geometry, exposure, sensor layout, and maintenance history. Environmental and operational variation can exceed the signal caused by early damage. Robustness should be evaluated through cross-structure, cross-site, cross-season, and prospective testing. Models should detect when input data departs from their validated domain and defer uncertain cases rather than produce unjustified confidence.
Domain adaptation seeks to reduce the mismatch between training and target structures. It can align features, update models using limited target data, or combine simulation with field observations. Recent unsupervised subdomain-adaptation research has addressed structural damage detection when labeled target-domain damage is unavailable [
78]. Such methods are promising, but alignment can also remove damage-sensitive differences if the domains are not defined carefully.
Out-of-distribution detection is complementary to adaptation. Instead of forcing every input into a known class, a model identifies when the observation falls outside its validated experience. For structural use, abstention should trigger a defined action: manual inspection, additional sensing, analytical review, or conservative classification. The abstention rate and downstream workload should be included in evaluation.
Beyond adaptation and out-of-distribution detection, recent studies have explored other ways to improve the reliability and interpretability of AI-enabled structural assessment. These include cloud-based digital-twin frameworks for continuous structural monitoring [
79] and expert-guided multitask learning for explainable damage assessment [
80].
6.4. Validation, Uncertainty, Explainability, and Human Oversight
Operational adoption requires integration with inspection records, sensor systems, BIM, analytical models, digital twins, and maintenance-management platforms. Common identifiers are needed to link observations to components and interventions over time. Data ownership, access, retention, cybersecurity, and model versioning must be governed. A secure audit trail should record the data, model, thresholds, uncertainty, and human approvals behind each recommendation.
Interoperability is partly semantic. The same term may describe a pixel-level crack class, an inspector-assigned defect, a component condition state, or a structural limit state. Shared ontologies and component identifiers are needed so that an observation can be traced through analysis and intervention. BIM can support geometry and identity, but it should not be assumed that an as-designed model represents the as-is building.
Cybersecurity is a structural-safety concern when monitoring data influences decisions. Spoofed sensors, altered records, unavailable communication, or unauthorized model updates can produce false alarms or conceal deterioration. Security requirements should include authentication, integrity checks, backups, privileged access, and safe degradation when connectivity is lost.
Structural assessment and rehabilitation occur within a regulated professional context. Codes and standards define loads, material requirements, acceptance criteria, investigation procedures, performance objectives, and professional responsibilities. AI output does not replace these requirements. Its permissible role depends on whether it supplies evidence, performs a calculation, recommends an action, or issues a safety-relevant conclusion.
For existing buildings, this obligation is especially important because the governing assessment basis is not identical to new design. Applicable building codes, existing-structure assessment standards, seismic evaluation and retrofit provisions, and material-specific repair requirements determine how observed deficiencies are investigated and accepted. AI may organize evidence or support calculations, but the engineer must still reconcile incomplete records, as-built deviations, deterioration, prior interventions, and the evidentiary limits permitted by the relevant jurisdiction and standard.
Current structural codes rarely prescribe validation procedures for learned models. This does not prevent their use, but it shifts responsibility to the engineer and organization to demonstrate fitness for purpose. A defensible assurance case should define the intended use, validated domain, data requirements, model version, performance criteria, uncertainty, human controls, fallback procedure, and monitoring plan. Claims should be no broader than the evidence supporting them.
Liability becomes difficult when several organizations contribute to the D2D Continuum. A sensor vendor may guarantee measurement performance, a software provider may supply a model, a consultant may interpret outputs, and an owner may authorize intervention. The interfaces among these roles should be contractual as well as technical. The record should identify who reviewed data quality, who accepted the diagnosis, who performed the structural assessment, and who approved the action.
Safety assurance should address systematic and random failures. Random errors are reflected in statistical performance and uncertainty. Systematic failure can arise from biased datasets, incorrect preprocessing, software defects, misunderstood units, inappropriate model reuse, or a change in the building. Conventional software verification remains necessary even when model performance is statistically evaluated. Units, coordinate systems, component identifiers, thresholds, and data transformations should be tested explicitly.
Change control is especially important. Retraining, updating a foundation model, changing a sensor, or modifying preprocessing can alter system behavior. A model that passed validation should not be silently replaced. Versioned deployment, regression testing, approval records, and rollback capability are needed. If online learning is used, the update rules and safeguards should be defined before deployment.
Professional review should include an abstention pathway. When data quality is poor, the case is out of distribution, model confidence is low, or alternative explanations are plausible, the system should recommend additional investigation rather than force a classification. Safe refusal is a capability, not a failure. Its frequency and operational cost should be measured because an overly cautious system may create an unsustainable inspection burden.
Regulatory acceptance will likely progress through bounded applications. Automated measurement of clearly visible defects, triage of inspection images, and anomaly screening are easier to validate than autonomous safety assessment or retrofit selection. Evidence collected in these bounded uses can support gradual expansion. The appropriate trajectory is therefore staged assurance: demonstrate value and control at one decision level before increasing autonomy or consequence.
6.5. Implementation Pathway and Technology Readiness
A practical D2D deployment can be organized into six controlled steps. First, the owner and engineer define the decision and the consequence of error. Second, they select observations that can inform that decision. Third, the AI model is validated on data representative of the building and acquisition process. Fourth, outputs are integrated with structural analysis, records, and uncertainty. Fifth, a human gate accepts, rejects, or requests additional evidence. Sixth, the outcome and subsequent observations are recorded.
This pattern discourages technology-first deployment. Purchasing a drone, installing sensors, or selecting a neural network before defining the decision often produces data without a governed use. Conversely, beginning with the decision clarifies required spatial resolution, monitoring duration, uncertainty, response time, and explanation.
Implementation should begin with shadow operation. The AI system processes cases, but established practice remains authoritative. Differences are reviewed and failure modes documented. Assisted operation can follow, in which the system prioritizes or summarizes evidence while an engineer performs the assessment. Higher levels of automation should be considered only where the use is bounded, monitoring is continuous, and safe fallback is available.
Performance monitoring should continue after deployment. Data distributions, abstention rates, false alarms, missed cases, calibration, processing time, and user overrides should be reviewed. Overrides are particularly informative: repeated disagreement may indicate model drift, a missing input, or a mismatch between the training target and the actual decision. These records turn professional judgment into evidence for system improvement without treating every past decision as unquestionable ground truth.
Economic evaluation should include the full workflow. Model development may represent a small portion of cost compared with data collection, labeling, sensor maintenance, software integration, cybersecurity, review, and retraining. Benefits may arise from reduced access requirements, earlier detection, consistent screening, or targeted intervention. A business case based only on faster inference is incomplete.
Finally, implementation should preserve the independence of safety review. If the same model proposes an intervention and estimates its benefit, shared bias may go undetected. Independent analytical checks, alternative models, inspections, or peer review provide defense in depth. AI is most valuable when it strengthens the evidence available to engineers, not when it removes the checks that make structural decisions trustworthy.
The evidence is most mature for automated recognition of visible defects and least mature for risk-informed rehabilitation selection. This imbalance should temper claims that AI already enables end-to-end autonomous management. The gap is not merely a lack of more accurate algorithms. It includes missing longitudinal datasets, weak cross-building validation, limited uncertainty propagation, incomplete connection to codes and mechanics, and few prospective studies of intervention outcomes.
The latest comprehensive reviews confirm continued growth in AI-based SHM while repeatedly identifying the same translational barriers [
34,
35]. This persistence suggests that the field needs a change in evaluation incentives. Publishing another architecture on a familiar benchmark contributes less to practice than careful cross-building validation, a documented failure analysis, or a study showing how uncertainty changes an engineering decision.
Reproducibility and delivery remain limited. A targeted availability check of the 33 application-oriented studies found only three that provided either a public code repository, a deployable application, or a software package at the time of review: the interpretable TRM-wall capacity model was deployed as a web application [
81], the MCO-beam study deposited data and code in a public GitHub repository [
82], and VDInM was published as a reusable life-cycle decision agent [
67]. Thus, approximately 9% (3/33) offered a directly accessible implementation. The count is conservative because supplementary files and repositories can change after publication, but it shows that reported model performance rarely translates into an independently executable tool for researchers or practitioners.
Evidence maturity should be assessed at the level of the complete use case rather than the algorithm alone. A crack detector may show strong performance under controlled imaging while the broader workflow for locating the defect on a building, verifying scale, updating the condition record, and triggering action remains less mature. The D2D evidence-maturity categories in
Figure 4 and
Table 4 are intended to make this system-level distinction explicit.
For the stage-level synthesis,
Figure 4 shows the distribution of the study-level evidence classifications reported in
Table S1. The values within each bar indicate the number of studies assigned to each evidence level. For example, for detection, 2 studies were classified as E1, 4 as E2, and 5 as E3, giving 11 study assignments for this stage. Studies covering more than one D2D stage contribute to each applicable stage, so the counts across stages are not mutually exclusive. Across the 33 application-oriented sources, none met all requirements for E4, including prospective or multi-site operation, cross-site or temporal validation, and uncertainty reporting, and none met the requirements for E5. The evidence across all eight D2D stages therefore remains within E1–E3 when the criteria in
Table 4 are applied consistently.
7. Research and Implementation Roadmap
Because the translation barriers identified throughout this review, and consolidated as evidence gaps in
Section 6.5, are interdependent, the roadmap follows the logic of the Data-to-Decision (D2D) Continuum to convert those gaps into staged research and implementation actions.
Figure 5 presents the time-phased progression, and
Table 5 maps each evidence gap to near-, medium-, and long-term actions.
7.1. Near-Term Priorities
Near-term work should improve data and reporting foundations. Priorities include building-specific multimodal datasets, explicit separation of training and test structures, reporting of negative and ambiguous cases, calibrated uncertainty, and shared definitions of D2D outputs. Studies should state whether they detect appearance, physical damage, performance loss, or decision urgency.
7.2. Medium-Term Priorities
Medium-term research should connect stages that are currently evaluated in isolation. Detection and localization outputs should be propagated into model updating and capacity assessment; prognosis should incorporate environmental exposure and intervention history; and risk models should use calibrated predictive distributions. Digital twins should increasingly be evaluated as governed information architecture rather than presented only as conceptual diagrams.
7.3. Long-Term Directions
Long-term systems may support adaptive monitoring, semi-autonomous inspection, and continuously updated intervention planning. Foundation models and AI agents may coordinate data retrieval, model execution, and reporting, but safety-critical actions must remain bounded by verified tools, explicit constraints, and human authorization. The target is not unrestricted autonomy; it is reliable assistance that reduces routine workload while making uncertainty and accountability more visible.
7.4. Stakeholder Actions
Researchers should prioritize transferable validation and publish failure cases. Practitioners should define decision requirements before selecting models and maintain human review at critical gates. Owners should invest in data continuity, component identifiers, and post-intervention records. Regulators and professional bodies should develop reporting expectations for data provenance, model validity, uncertainty, and responsibility. Collaboration among these groups is necessary because no single stakeholder controls the complete D2D Continuum.
7.5. Sustainable and Resilient Rehabilitation
AI can support sustainability when it helps avoid premature replacement, target intervention, extend service life, and compare safety benefits with cost, carbon, and disruption. Sustainability claims should include rebound and uncertainty effects: additional sensing and computation have impacts, and an incorrect life-extension recommendation can defer necessary work. A mature D2D Continuum should therefore support multi-criteria decisions without reducing safety to one objective among many.
7.6. Priority Research Questions
The questions below are organized according to the stages of the D2D Continuum so that future research priorities remain aligned with engineering decisions rather than algorithm development alone. The next generation of research should be guided by decision-relevant questions rather than model novelty alone.
For data acquisition, the central issue is determining the minimum combination of sensing, metadata, and quality control required to support a defined assessment objective. This includes establishing when low-cost imagery is sufficient, when multimodal sensing materially reduces uncertainty, and how sensor-placement decisions influence the detectability of plausible damage scenarios.
For detection and localization, the key question is whether model performance transfers to buildings with different materials, ages, geometries, surface conditions, and acquisition environments. Cross-building evaluation should become a standard requirement. Studies should report which defects are missing, how models respond to unfamiliar inputs, and how much expert review remains necessary. Improvements on established benchmarks should not be interpreted as evidence of field readiness without external validation.
For quantification, AI outputs must be connected to traceable physical quantities. Crack geometry, stiffness loss, corrosion state, residual deformation, and damage indices require different calibration procedures and uncertainty models. Evaluation should therefore identify the smallest prediction error capable of changing an engineering decision and test models against that threshold. Such decision-sensitive evaluation is more informative than reporting average error without considering its structural significance.
For condition and performance assessment, the principal challenge is reconciling data-driven inference with structural mechanics and code-based evaluation. Hybrid approaches should be tested under both accurate and incomplete physical models. Direct comparisons among purely data-driven, physics-guided, and conventional analytical methods should use the same evidence and explicitly report cases in which their conclusions disagree. These disagreements may reveal limitations in the data, the physical assumptions, or both.
For prognosis, longitudinal building datasets remain essential. Research should examine how deterioration, inspection, repair, occupancy, and environmental exposure interact over time. Remaining-useful-life models should clearly define the end state, account for censoring, and provide calibrated predictive distributions rather than isolated point estimates. They should also be evaluated under changes in inspection frequency and intervention policy, since these practices affect both deterioration and the observations from which prognosis is learned.
For risk assessment, model uncertainty must be propagated into probability and consequence estimates rather than appended as an isolated confidence value. Decision metrics should reflect asymmetric errors and the possibility of rare but high-consequence states. The practical question is whether an AI-assisted assessment changes risk estimates, inspection priorities, or intervention decisions sufficiently to justify its additional complexity.
For rehabilitation, the immediate need is a shared and auditable representation of intervention alternatives, eligibility constraints, expected structural effects, durability, cost, carbon, disruption, and uncertainty. Prospective case studies should compare AI-supported recommendations with multidisciplinary engineering decisions and document post-intervention performance. Without outcome-based evidence, rehabilitation models will continue to optimize simulated policies without demonstrating whether their recommendations succeed in practice.
Across all stages of the D2D Continuum, a decisive research question is when a model should abstain from producing an output. Abstention, escalation, and requests for additional evidence should be designed and evaluated as explicit system capabilities. Ultimately, a trustworthy D2D Continuum is not one that always produces a result; it is one that recognizes when the available evidence and validated knowledge are insufficient for the consequence of the decision.
Taken together, these stage-specific questions define a decision-oriented research agenda and are consolidated in
Table 5, which maps the identified evidence gaps to near-, medium-, and long-term research and implementation priorities.
8. Conclusions
What has AI already achieved? The evidence reviewed here shows mature or rapidly maturing capability for bounded acquisition, defect screening, image-based detection, spatial localization, and anomaly recognition. AI can process inspection and monitoring data at scales that are difficult to sustain manually, and it can support consistent triage when the target, acquisition conditions, and validated domain are explicit. A smaller but important body of work also demonstrates decision-relevant links to damage quantification, residual performance, deterioration prognosis, reliability updating, and maintenance optimization. These achievements are substantive, but they are stage-specific; they do not yet constitute a validated autonomous assessment-and-rehabilitation system.
What remains unresolved? The principal gap is continuity of evidence across the engineering decision chain. Detection outputs are not consistently registered to structural components, quantified damage is not always connected to mechanics or limit states, prognostic uncertainty is rarely propagated into risk, and rehabilitation recommendations are seldom evaluated through implemented outcomes. Specifically, cross-building transfer, representative longitudinal datasets, calibrated uncertainty, trustworthy abstention, semantic interoperability, and code-aligned assurance remain uneven. The reviewed literature also reveals a pronounced maturity gradient: evidence is strongest at the upstream diagnostic stages and weakest where consequences and professional responsibility are greatest.
What should researchers and practitioners do next? Research should evaluate complete use cases rather than isolated model accuracy, use structure- and event-level validation, report failure and out-of-domain cases, and preserve engineering units and uncertainty at every interface. Building specific evidence should be combined with population learning and physical models, while prospective studies should record whether interventions were feasible, selected, implemented, and effective. Practitioners should introduce AI through bounded shadow and assisted-operation workflows with version control, audit trails, explicit human gates, and independent safety checks.
The proposed Data-to-Decision (D2D) Continuum provides the organizing contribution for this agenda. It reframes AI-enabled structural assessment as a sequence of accountable transformations from traceable observations to authorized action, followed by post-intervention learning. Its purpose is not to prescribe one algorithm or platform, but to make visible the output contract, uncertainty, validation claim, and responsible decision-maker at every stage. The defining vision is therefore not autonomous AI that replaces structural judgment, but an evidence-connected engineering system in which every observation can be traced to a safer, more transparent, and more sustainable engineering decision.