Next Article in Journal
Impact of Regulation of Wax-Based and Bio-Oil-Based Warm-Mix Additives on the Phase Behavior and Rheological Properties of Rubber-Modified Asphalt
Previous Article in Journal
Clay and Microsilica Additives’ Effect on the Properties and Structure of Injectable Cement–Clay Mortars for Soil Consolidation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Machine Learning for Structural Steels: Materials Design, Property Prediction, Durability, and Future Directions

1
School of Mechanical and Civil Engineering, Jilin Agricultural Science and Technology College, Jilin 132101, China
2
Faculty of Engineering, University Malaysia Sabah, Jalan UMS, Kota Kinabalu 88400, Sabah, Malaysia
*
Author to whom correspondence should be addressed.
Materials 2026, 19(17), 3612; https://doi.org/10.3390/ma19173612
Submission received: 19 July 2026 / Revised: 11 August 2026 / Accepted: 20 August 2026 / Published: 25 August 2026

Abstract

Machine learning (ML) provides new opportunities to model the nonlinear relationships among composition, processing, microstructure, defects, properties, and in-service degradation of structural steels. This structured critical review examines ML applications to materials and process design, microstructural characterization, mechanical-property prediction, corrosion, fire and elevated-temperature performance, fatigue, fracture, and remaining-life assessment. Literature published up to 31 July 2026 was searched primarily through the Web of Science Core Collection and Scopus. A total of 110 publications were retained based on their relevance to structural steels, transparency of data and modeling procedures, and availability of information on validation or engineering applicability. The reviewed studies show that model suitability depends strongly on data modality, sample independence, feature representation, and validation strategy rather than on algorithm family alone. ML has progressed from property prediction toward process optimization, inverse materials design, environmental degradation assessment, and fatigue- and crack-related prognostics. However, independent cross-manufacturer, cross-laboratory, production-scale, and field validation remains limited, while uncertainty quantification and applicability-domain assessment are still inconsistently reported. These limitations are particularly important for corrosion, fire, fatigue, and remaining-life applications, where internally validated models should not be interpreted as substitutes for established physical models or design provisions. Future research should prioritize standardized multimodal data, physics-informed and uncertainty-aware modeling, prospective validation, and rigorously evaluated closed-loop monitoring and digital-twin frameworks for structural-steel life-cycle management.

1. Introduction

Structural steel combines high load-carrying efficiency, favorable ductility and toughness, mature fabrication and welding technologies, and substantial recyclability, making it one of the principal load-bearing materials in buildings, bridges, industrial facilities, long-span structures, offshore engineering, and transportation infrastructure. As structural systems move toward greater height, longer spans, increased prefabrication, and extended service life, steel design is no longer governed by strength alone but increasingly requires the simultaneous control of strength, ductility, toughness, weldability, corrosion resistance, fire resistance, fatigue resistance, and cost. These competing requirements have made the design and assessment of structural steels a multi-variable and multi-objective problem. In this review, structural steels primarily refer to conventional carbon structural steels, high-strength low-alloy steels, weathering steels, fire-resistant steels, bridge and pipeline steels, structural stainless steels, and their welded products used in civil infrastructure applications.
The engineering properties of structural steels emerge from coupled relationships among chemical composition, manufacturing history, microstructure, defects, and service conditions. Alloying, steel cleanliness, continuous casting, rolling and cooling schedules, heat treatment, and welding thermal cycles jointly determine phase constitution, grain structure, precipitation, dislocation substructure, texture, inclusions, and local defects, which in turn control strength, ductility, toughness, fatigue resistance, and environmental degradation [1,2,3]. Conventional approaches based on empirical alloy development, physical metallurgy, thermodynamic and kinetic calculations, and numerical simulation provide important mechanistic insight, but become increasingly expensive or restrictive when applied to high-dimensional composition spaces, complex processing routes, multiple performance objectives, and coupled service environments [4,5,6,7,8]. Machine learning (ML) provides a complementary route by learning nonlinear relationships from experimental, computational, industrial, imaging, and monitoring data while retaining the possibility of incorporating physical knowledge.
Applications of ML in steels have consequently expanded well beyond conventional property regression. Current studies encompass phase-transformation prediction, process optimization, microstructure classification and segmentation, defect detection, mechanical-property prediction, multi-objective alloy design, inverse microstructure design, corrosion modeling, elevated-temperature assessment, fatigue and fracture prediction, and remaining-life evaluation [3,9,10,11,12,13,14,15,16,17]. This expansion is accompanied by increasing diversity in data modality, ranging from tabular composition–process datasets and microstructural images to three-dimensional defect data, time-series sensor signals, and physics-based simulations. Accordingly, model performance cannot be judged solely by the choice of algorithm or by a single accuracy metric. The relevance of the input representation, independence of the validation data, physical consistency of predictions, and evidence for transfer beyond the original dataset are equally important for engineering applications.
Existing reviews, however, remain fragmented across these topics. Reviews of ML for alloys and metallic materials have mainly emphasized descriptors, property prediction, microstructure analysis, alloy design, and inverse optimization [4,5,6], whereas specialized reviews have focused on steel microstructures, non-metallic inclusions, constitutive behavior, or fatigue prediction [2,13,14,15,16]. Reviews in construction materials have predominantly addressed concrete and other building materials [17], while AI-oriented reviews of steel structures have concentrated on beams, columns, joints, connections, structural capacity, seismic response, and structural health monitoring rather than on the metallurgical origin of material performance [18]. Reviews closer to steel manufacturing have provided valuable syntheses of process modeling and performance prediction but generally do not connect material design with environmental degradation and service-life assessment [19,20,21,22]. The distinction between these previous perspectives and the present review is summarized in Table 1.
Accordingly, the novelty of this review lies not in cataloguing machine-learning algorithms across steel systems, but in integrating materials design, service degradation, and evidence-strength assessment within a unified materials-to-service framework for civil-infrastructure structural steels.
To maintain a clear scope while still capturing transferable advances from adjacent fields, the evidence considered in this review is classified into three levels. Direct evidence refers to studies on structural steels and welded steel products used in buildings, bridges, pipelines, offshore structures, and related civil infrastructure. Adjacent steel-alloy evidence includes studies on related engineering steels whose metallurgical mechanisms or modeling strategies are relevant to structural-steel applications. Methodological transfer cases include studies on materials such as reduced-activation ferritic–martensitic steels, twinning-induced plasticity steels, nuclear steels, reinforcing steel in cementitious systems, and other non-structural alloy systems and are included only when they demonstrate methods that can reasonably be transferred to structural-steel research. A second distinction is made between prediction scales. The primary scope of the review is the material scale, whereas specimen/joint-, member-, and structural-system studies are included only when they explicitly propagate material degradation, material-state variables, or material-model outputs into engineering performance or life-cycle assessment. This distinction is particularly important in fire, fatigue, and durability applications, where material degradation may propagate across multiple structural scales.
On this basis, the review is organized around three questions: (RQ1) What ML tasks have been investigated across the design, manufacturing, characterization, and service stages of structural steels? (RQ2) What data modalities, feature representations, and validation strategies have been used, and to what extent do they support generalization beyond the original datasets? (RQ3) Which approaches have demonstrated meaningful engineering validation, and what barriers currently prevent the deployment of ML in structural-steel design and life-cycle management? To address these questions, the literature is synthesized according to material relevance, data modality, modeling objective, validation strategy, and engineering evidence, while distinguishing direct structural-steel studies from adjacent evidence and methodological transfer cases. Section 2 describes the review methodology. Section 3 establishes the principal data sources, machine-learning workflow, and trustworthy evaluation criteria. Section 4 reviews composition and process design, microstructure modeling, mechanical-property prediction, and inverse materials design. Section 5 examines corrosion, fire and extreme-temperature performance, fatigue, fracture, and remaining-life assessment. Section 6 synthesizes the current challenges and future research directions, and Section 7 presents the principal conclusions.

2. Review Methodology

A structured literature search was conducted to identify studies on machine learning (ML) relevant to the design, manufacturing, characterization, property prediction, durability, and service-life assessment of structural steels. The Web of Science Core Collection and Scopus were used as the primary bibliographic databases, supplemented by targeted searches and backward and forward citation tracking of relevant reviews and primary studies. The literature search covered publications available up to 31 July 2026, with no lower publication-year restriction. Only English-language publications containing sufficient information on the material system, dataset, modeling procedure, or validation strategy were considered. The search combined terms related to computational methods, steel materials, and engineering tasks, including machine learning, deep learning, artificial intelligence, data-driven, structural steel, high-strength steel, weathering steel, pipeline steel, microstructure, property prediction, corrosion, fire, fatigue, fracture, and remaining life.
To avoid interpreting all reported prediction accuracies as equivalent evidence, validation strength was classified into four levels: V1, internal random train–test partitioning or conventional cross-validation; V2, group-based validation using independent heats, specimens, literature sources, or production periods; V3, external validation using an independent laboratory, production line, manufacturer, steel grade, or separately produced heat; and V4, engineering-scale validation involving industrial production, full-scale testing, field monitoring, or operational infrastructure. Because the reviewed studies differ substantially in materials, target variables, dataset structures, and performance metrics, their reported accuracies were not statistically pooled. Instead, the evidence was critically synthesized according to material relevance, data modality, physical representation, validation strength, uncertainty treatment, and experimental or engineering verification. This framework was used consistently in the subsequent comparison of ML applications to structural-steel design, property prediction, durability, and service-life assessment.

3. Fundamentals of Structural Steels and Machine Learning

3.1. Material Characteristics and Data Sources of Structural Steels

Machine-learning data for structural steels are primarily derived from experimental testing, published literature and open databases, industrial production, materials characterization, in-service monitoring, and numerical simulation. These data sources differ not only in volume and modality but, more importantly, in their degree of statistical independence and physical fidelity. A dataset containing thousands of records does not necessarily contain thousands of independent material states. Repeated specimens from one heat, cropped patches from the same micrograph, dense sensor measurements from one production campaign, and repeated measurements under the same exposure condition may be strongly correlated. Dataset size should therefore be reported together with the number of independent heats, specimens, exposure conditions, or production campaigns rather than using record count alone as an indicator of information content [19,20,21].
Experimental data provide the most direct link to material behavior but are costly to generate for fracture toughness, long-term corrosion, creep, very-high-cycle fatigue, and other time- or resource-intensive properties. Literature-derived datasets can broaden the range of steel grades, processing conditions, and service environments, but introduce heterogeneity associated with grade nomenclature, specimen geometry, sampling orientation, testing temperature, loading rate, test standards, and reporting precision. Industrial datasets may contain very large numbers of process records but considerably fewer independent heats or production campaigns. They are also affected by sensor drift, equipment maintenance, process changes, abnormal operating conditions, and manual-entry errors [19,20,21,22]. Consequently, the effective information content of an industrial dataset depends on both the number of records and the diversity of independent material and process conditions represented.
Steel production generates tabular and time-series variables including chemical composition, temperature, processing time, flow rate, pressure, rolling force, roll gap, and cooling-water flow rate. These data are particularly suitable for endpoint prediction, process optimization, anomaly detection, and quality early warning. However, randomly separating individual records collected from the same heat or production period may result in information leakage. Industrial datasets should therefore be partitioned by heat, production campaign, or chronological production period whenever the intended application involves prediction for future production or previously unseen material batches.
Microstructural and defect data provide an intermediate representation between manufacturing history and macroscopic performance. Optical microscopy (OM), scanning electron microscopy (SEM), electron backscatter diffraction (EBSD), and transmission electron microscopy (TEM) characterize phase constitution, grain structure, crystallographic texture, precipitates, and inclusions, whereas X-ray computed tomography (X-CT) can provide three-dimensional information on internal pores, inclusions, and cracks. For image-based machine learning, the number of cropped patches should be distinguished explicitly from the number of original specimens or independently acquired fields of view. Randomly allocating patches from the same parent image to both training and test sets can lead to severe data leakage and substantially overestimate generalization.
Several open datasets illustrate both the opportunities and limitations of image-based materials learning. The Aachen–Heerlen dataset contains 1705 SEM images and 8909 expert-annotated polygons describing martensite–austenite islands [23]. Its article and associated dataset are persistently identifiable through DOI-based records, facilitating reproducible reuse. The Ultrahigh Carbon Steel Micrograph Database (UHCSDB) links micrographs acquired at different length scales with heat-treatment and microstructural metadata [24]. The ferritic-steel X-CT database provides three-dimensional defect information for mechanically deformed specimens and is accompanied by a persistent Dryad repository record [25]. These resources demonstrate that high-quality labels, specimen-level identifiers, imaging conditions, and access to original rather than only cropped images are critical for reliable reuse.
Data imbalance and censoring require task-specific treatment. Rare brittle-fracture events, severe localized corrosion, extreme inclusions, and very short fatigue lives may constitute a small fraction of a dataset but can dominate structural reliability. They should not be discarded simply because they are statistical outliers. Similarly, fatigue run-outs, particularly in high- and very-high-cycle regimes, are censored observations rather than missing values. Treating run-outs as ordinary failures or removing them from the dataset can bias life predictions. Depending on the task, survival-analysis concepts, censored regression, probabilistic life models, or imbalance-aware sampling should therefore be considered instead of conventional imputation or unrestricted resampling.
Microstructural representation must also retain physically relevant morphology and spatial information. Conventional computer-vision and microstructure-reconstruction studies have demonstrated that phase fraction alone cannot uniquely characterize complex microstructures; morphology, spatial statistics, connectivity, characteristic length scales, and interfacial distributions may contain additional information relevant to mechanical response [26,27]. High-quality labels should therefore reflect metallurgical definitions rather than image contrast alone, and the imaging scale, pixel size, specimen preparation, acquisition conditions, and annotation procedure should be preserved as part of the dataset metadata.
Numerical data are mainly generated through the CALculation of PHAse Diagrams (CALPHAD) method, density functional theory (DFT), molecular dynamics, phase-field modeling, and finite element analysis. Such approaches can efficiently expand the explored composition and processing spaces and provide information on phase stability, microstructural evolution, local stress, and damage variables that may be difficult to measure experimentally. However, computational records should not be pooled indiscriminately with experimental observations. Each numerical dataset should retain an explicit fidelity label and information on governing assumptions, constitutive parameters, boundary conditions, calibration data, and numerical uncertainty.
The discrepancy between simulations and experiments should itself be treated as a source of uncertainty. High- or low-fidelity numerical data can be used for pretraining, surrogate-model development, or multi-fidelity learning, followed by calibration against independent experimental measurements [28]. Uncertainty originating from input parameters, constitutive models, boundary conditions, or numerical approximations should, where possible, be propagated into the final machine-learning prediction. In this context, computational data are most valuable as a source of physical structure and additional coverage rather than as an unquestioned substitute for experimental ground truth. Elemental-composition networks and crystal-graph models likewise demonstrate the potential for automatic representation learning from atomic-scale information, but their application to engineering structural steels still requires scale bridging through processing history, microstructure, defects, and service-state variables [10,11].
As summarized in Table 2, different data modalities provide complementary information. Composition and processing records are efficient for property prediction and process optimization but describe local heterogeneity only indirectly. Microstructural and three-dimensional defect datasets provide stronger links to physical mechanisms but are more expensive to acquire and annotate. Industrial sensors offer high temporal density but may contain relatively few independent production conditions. Simulation can efficiently expand the design space but introduces model-form and parameter uncertainty. Machine learning for structural steels should therefore progress from indiscriminate aggregation of heterogeneous records toward multimodal datasets in which provenance, independence, fidelity, and uncertainty are explicitly represented.
Figure 1 follows the materials-science sequence of composition, manufacturing process, microstructure, defect state, macroscopic properties, and in-service degradation. Experimental data, industrial data, microstructural images, sensor signals, and simulation results are linked to property prediction, process optimization, defect identification, durability assessment, and inverse materials design.
Data traceability is essential for reproducibility and model transfer. Structural-steel datasets should, at a minimum, retain the steel grade, heat number, product form, plate or section thickness, sampling location and orientation, manufacturing route, welding procedure, heat-treatment history, specimen geometry, testing standard, loading or strain rate, testing temperature, service environment, characterization equipment, calibration state, preprocessing procedure, and measurement uncertainty. Both data and metadata should conform to the Findable, Accessible, Interoperable, and Reusable (FAIR) principles [29] and maintain a traceable lineage from the original experiment or simulation to the variables used for model training [30]. Without such information, a model may inadvertently interpret batch effects, laboratory practices, equipment differences, or data-processing choices as intrinsic materials relationships.

3.2. Machine-Learning Workflow and Model Evaluation

Machine-learning studies of structural steels typically involve problem formulation, data and metadata curation, preprocessing, physics-based feature construction, model training, validation, and engineering deployment. Research in materials informatics has shown that model performance is highly sensitive to descriptor selection, data partitioning, and hyperparameter-optimization procedures. Comparisons among algorithms are therefore meaningful only when they are conducted using consistent datasets and validation protocols [7,8,9,31,32,33,34,35].
To ensure a consistent evaluation framework across different structural-steel applications, this review summarizes machine-learning workflows from data partitioning, model development, performance evaluation, uncertainty analysis, and engineering validation perspectives. These criteria are subsequently used to assess the reliability of reported studies in Section 3 and Section 4.
Figure 2 presents a sequential workflow comprising problem definition, data acquisition and metadata recording, data cleaning, physics-based feature construction, model training and hyperparameter optimization, group-based validation, interpretability analysis, domain-of-applicability assessment, uncertainty quantification, and independent external validation. Particular emphasis should be placed on grouping samples according to heat, steel grade, manufacturer, or production period to prevent data leakage caused by conventional random partitioning.
Data preprocessing requires the standardization of steel-grade designations, composition units, processing variables, and property definitions, together with systematic examination of missing values, duplicate records, and outliers. Anomalous observations should not be removed solely on the basis of statistical thresholds, because low-temperature brittle fracture, exceptionally short fatigue life, or severe corrosion may represent the most engineering-relevant boundary cases. For variables recorded only for specific steel grades or processing conditions, simple mean imputation may also generate artificial material states that do not physically exist.
The purpose of feature construction is not to increase the number of variables, but to establish effective representations of the material state. In addition to raw chemical compositions, physically meaningful variables such as carbon equivalent, hardenability indices, cooling rate, phase fraction, grain size, inclusion size, and defect density may be introduced. General-purpose materials machine-learning frameworks have shown that descriptors derived from elemental properties can increase the information density of composition-based datasets [32]. Open-source tools such as Matminer can further standardize descriptor generation and unify data interfaces [33]. However, large numbers of automatically generated variables may introduce redundancy and multicollinearity. Feature selection must therefore be conducted exclusively within the training data and should not use the complete dataset before the test set has been separated.
The model type should be matched to the sample size and data modality. Random forest, support vector machines, and gradient-boosting models are generally suitable for small- to medium-sized tabular datasets. Convolutional neural networks are appropriate for microstructural and defect images, whereas recurrent neural networks (RNNs) and Transformer architectures are suitable for industrial time-series signals. The principal advantage of deep learning lies in its ability to learn hierarchical representations automatically, rather than simply increasing the number of hidden layers [9,12]. For small-data problems in structural-steel research, physics-based features, regularization, transfer learning, and group-based validation generally improve generalization more effectively than increasing network depth [34].
Data partitioning is among the most frequently overlooked aspects of model evaluation. Different specimens from the same heat, cropped regions from the same image, and repeated tests reported in the same publication are often strongly correlated. If such samples are distributed across both the training and test sets, model performance can be substantially overestimated. Composition–property datasets should therefore be grouped by heat or steel grade, industrial datasets should be divided chronologically, literature-derived datasets should be grouped by source, and microstructural images should be partitioned according to the original specimen. Materials benchmarking studies have demonstrated that differences in data cleaning, partitioning strategies, and hyperparameter optimization can introduce model-selection bias. Model performance is comparable only when consistent datasets and validation procedures are used [35].
In materials informatics, the reported prediction accuracy strongly depends on how training and testing data are separated. Random splitting may lead to overly optimistic estimates when samples from the same heat, production campaign, specimen series, or image source are distributed into both subsets. Therefore, group-based partitioning according to steel grade, heat number, laboratory, production line, or exposure condition is preferred for evaluating engineering generalization. Model performance should be distinguished among three prediction scenarios: interpolation within the known design space, extrapolation within related steel families, and out-of-distribution prediction beyond the original data domain. These scenarios represent different levels of engineering confidence and should not be interpreted using a single test-set metric.
Regression tasks are commonly evaluated using the coefficient of determination (R2), mean absolute error (MAE), and root mean square error (RMSE). Classification tasks typically use accuracy, precision, recall, the F1 score, and the area under the receiver operating characteristic curve (AUC). Average metrics alone are insufficient for safety-critical applications. A model may exhibit low overall error within the conventional data range while systematically underestimating brittle-fracture risk, short fatigue lives, or high-temperature instability. Region-specific errors, the proportion of non-conservative predictions, prediction-interval coverage, and calibration error should therefore also be reported.
The selection of evaluation metrics should be consistent with the engineering consequence of prediction errors. For regression tasks involving strength, toughness, corrosion rate, and fatigue life, commonly used indicators include R2, RMSE, MAE, and calibration-related measures. However, average errors alone cannot represent safety implications. Underprediction of strength and overprediction of service life may introduce non-conservative risks; therefore, direction-sensitive errors should also be considered. For classification tasks such as failure mode identification or risk assessment, accuracy alone is insufficient, particularly for imbalanced datasets. Balanced accuracy, Matthews correlation coefficient, precision–recall curves, area under the precision–recall curve, and class-specific recall should be reported when appropriate.
High test-set accuracy does not necessarily indicate reliable extrapolation. Domain-of-applicability analysis can identify the composition, microstructure, and processing regions within which a model is expected to maintain relatively low prediction errors [36]. When distribution shifts exist between the training data and a new steel grade or production line, performance obtained from a randomly selected test set may substantially overestimate real-world generalization [37]. Accordingly, leave-one-heat-out, leave-one-grade-out, cross-manufacturer, and cross-laboratory validation provide more meaningful assessments of engineering applicability than merely increasing the size of a random test subset.
When hyperparameter optimization or model selection is performed, nested cross-validation is recommended to avoid information leakage between optimization and evaluation procedures. Otherwise, even group-based validation may retain optimistic bias because model configurations have already been indirectly optimized toward the validation data.
Uncertainty quantification helps determine whether an individual prediction is sufficiently reliable for practical use. Prediction intervals, Gaussian-process models, ensemble methods, and quantile regression can provide risk warnings when an input approaches or exceeds the range represented by the training data [38]. Explainable machine learning can further clarify the contributions of composition, processing, and microstructural variables to model predictions. However, feature-importance measures and SHapley Additive exPlanations (SHAP) values describe statistical associations and do not, by themselves, establish causal mechanisms. Their interpretation must therefore be examined against theories of phase transformation, strengthening, fracture, and corrosion [39]. On this basis, machine-learning models for structural steels should be evaluated comprehensively across the six dimensions summarized in Table 3.
Feature importance methods such as SHAP and permutation importance provide information on model-specific associations rather than direct causal mechanisms. Therefore, interpretation results should be examined for stability across data partitions, model families, and perturbation analyses, and should be further supported by physical understanding or experimental evidence. In addition, uncertainty quantification is essential for safety-critical structural-steel applications. Prediction intervals, calibration analysis, applicability-domain assessment, and out-of-distribution detection should be considered to identify cases where models may produce unreliable predictions.
Overall, the reliability of machine-learning models for structural steels depends first on whether the available data adequately describe the material state, second on the data-partitioning and validation strategies, and only then on the choice of algorithm. High-quality studies should report not only predictive accuracy but also the applicable ranges of steel grades, processing conditions, and service environments, together with physical interpretation, uncertainty estimates, and independent validation results. Only when data are traceable, models are interpretable, and predictions are independently verifiable can machine learning progress from a property-fitting tool to a reliable method for structural-steel design and engineering decision-making.

4. Machine Learning Applications in Structural-Steel Design and Property Prediction

The applications discussed in this section are categorized according to the material design and evaluation hierarchy rather than algorithm type. First, machine learning is applied to composition, processing, phase transformation, and microstructure design, where the objective is to establish relationships between manufacturing history and internal structure. Subsequently, models are used for mechanical-property prediction and multi-objective inverse design, where predicted properties are transformed into material-selection strategies. Throughout this section, studies are evaluated not only by predictive accuracy but also by dataset independence, validation strategy, uncertainty treatment, and applicability domain.
The general workflow of machine learning for structural steels involves developing forward models from composition, manufacturing-process, and microstructural data, followed by the integration of physical-metallurgy knowledge and optimization algorithms to infer material compositions, processing schedules, or microstructures that satisfy prescribed performance targets. Compared with studies focused solely on predictive accuracy, data-driven structural-steel design places greater emphasis on the physical plausibility, manufacturability, and experimental verifiability of the predicted solutions.

4.1. Composition, Process, Phase Transformation, and Microstructure Design

The studies discussed in this section are classified according to their relevance to civil-infrastructure structural steels. Direct evidence refers to studies using steels explicitly developed or evaluated for structural applications in buildings, bridges, pipelines, or other civil infrastructure. Adjacent-steel evidence includes closely related high-strength, low-alloy, pipeline, or welding steels whose composition–process–microstructure relationships are relevant to structural-steel design. Methodological transfer cases refer to studies on other steel families, such as high-alloy, stainless, or automotive steels, which are included only when their machine-learning strategy provides transferable methodological insight. This distinction is maintained throughout the following discussion to avoid treating methodological similarity as evidence of direct engineering applicability.
Physics-guided modeling provides an important methodological transfer case. Shen et al. [40] incorporated precipitation-related physical descriptors into support-vector and classification models coupled with a genetic algorithm for ultrahigh-strength stainless-steel design. The physical descriptors helped exclude thermodynamically infeasible candidates, although the material system lies outside conventional civil-infrastructure structural steels.
To address the long experimental cycles required to determine time–temperature–transformation (TTT) and continuous-cooling-transformation (CCT) diagrams, Huang et al. [41] combined classification and ensemble regression to predict TTT diagrams for high-alloy steels. The dataset contained 58 TTT diagrams, with five complete steel curves held out for testing, thereby avoiding the direct mixing of the test-steel curves with the modeling data. However, these high-alloy steels are not representative of conventional civil-infrastructure structural steels and are therefore considered a methodological transfer case. Geng et al. developed a hybrid model for synthetic weld heat-affected-zone CCT diagrams of low-alloy steels using 97 complete CCT datasets, with 91 groups for model development and six steels reserved for testing [42]. Because the target was the welding heat-affected zone rather than bulk structural-steel transformation behavior, this study is considered adjacent-steel evidence. Similar hardenability models were developed for non-boron and boron steels using Jominy data [43,44]. These studies demonstrate the value of curve-level prediction for heat-treatment and alloy-design tasks, but their applicability to civil-infrastructure structural steels remains dependent on compositional overlap and validation using independent structural-steel grades.
Recent studies have further integrated transformation-product type, phase fraction, and microstructural morphology within unified modeling frameworks. Cao et al. first transformed reheating and deformation parameters into physically meaningful state variables, such as prior-austenite grain size and stored deformation energy. Gradient-boosted decision trees (GBDTs) and support vector machines were then used to predict the formation and volume fractions of polygonal ferrite, granular bainite, acicular ferrite, and lath bainite. The resulting models were subsequently applied to alloy-lean design and rolling–cooling route optimization for high-strength low-alloy steels. The predicted compositions and processing schedules were subsequently evaluated through newly produced steel and microstructural and mechanical-property characterization [45]. Because the study reports experimental validation of newly designed compositions but does not provide sufficient information to quantify the compositional distance from the training domain or to establish prospective validation in the strict sense, it is treated here as supporting external experimental evidence rather than definitive industrial-scale prospective validation.
This distinction is important because experimental confirmation of a model-generated candidate does not necessarily demonstrate generalization to unseen heats, manufacturers, or production campaigns. In the present review, industrial validation is considered strongest when newly produced heats are independent of model development and their selection is prospectively defined; otherwise, the evidence is classified as external experimental verification but not full production-line validation.
As illustrated in Figure 3, this class of approach does not directly predict final properties from processing parameters. Instead, processing variables are first converted into physical descriptors of the transformation state, after which microstructure type, phase fraction, and macroscopic properties are predicted sequentially. The resulting information is then fed back into composition and process design, forming a closed-loop optimization framework.
Microstructural characterization has evolved from manually engineered descriptors toward automatic feature learning. The conventional bag-of-visual-words approach can retrieve and classify microstructures without requiring a predefined feature catalogue, but it remains sensitive to imaging scale, specimen-preparation differences, and local texture variations [26]. Deep neural networks have improved the classification of martensite, tempered martensite, bainite, and pearlite in steels [46]. Nevertheless, image-based performance should be evaluated at the level of original specimens rather than cropped patches, because multiple patches obtained from one specimen are not statistically independent observations. Reliable segmentation also depends on the metallurgical validity and uncertainty of the labels. Studies using electron backscatter diffraction (EBSD) phase maps as supervisory labels can distinguish morphologically similar constituents through crystallographic orientation, grain-boundary characteristics, and intragranular misorientation [47]. However, label uncertainty arising from phase-definition criteria, annotation disagreement, and EBSD indexing quality should be distinguished from model segmentation error. Moreover, high segmentation accuracy does not necessarily imply equally accurate downstream prediction of mechanical properties; these two levels of performance should therefore be evaluated separately. A complementary strategy is to improve segmentation robustness through microscopy-specific pretraining. Stuckner et al. [48] pretrained deep-learning encoders on a large microscopy dataset and transferred the learned representations to microstructure-segmentation tasks, demonstrating improved performance relative to ImageNet-based pretraining, particularly when only limited task-specific training data were available. Although demonstrated using Ni-based superalloys rather than structural steels, this study provides a relevant methodological transfer case for data-limited microstructural analysis.
Figure 4 illustrates this microscopy-specific pretraining strategy using representative Ni-based superalloy segmentation results reported by Stuckner et al. [48]. The comparison shows that MicroNet-pretrained models preserve precipitate boundaries more reliably and retain greater sensitivity to fine tertiary precipitates as the amount of task-specific training data decreases, whereas ImageNet-pretrained models exhibit more pronounced over-segmentation and missed detections under few-shot and one-shot conditions. These results highlight the importance of domain-relevant pretraining for microstructure segmentation when experimentally annotated data are scarce. More broadly, however, segmentation accuracy should still be interpreted together with label quality, specimen-level independence, and the accuracy of subsequent quantitative microstructural measurements.
Table 4 summarizes representative applications of machine learning to the composition, processing, and microstructure design of structural steels. The ultimate purpose of microstructure recognition is not merely to produce visually refined segmentation maps, but to establish more accurate microstructure–property relationships. Liu et al. coupled deep-learning-based microstructure segmentation with a random-forest property model and used refined descriptors, including grain size, phase fraction, grain-boundary length, and second-phase interfacial length, to predict the properties of hot-rolled steels. Compared with models based only on average grain size and phase fraction, the RMSE values for yield strength, ultimate tensile strength, and elongation were reduced by 71.53%, 67.51%, and 53.40%, respectively [49]. These results demonstrate that morphological, interfacial, and spatial-distribution information can complement conventional average statistics and improve the representation of strengthening and plastic-deformation mechanisms.
Overall, the available studies can be classified into three technical routes. The first directly predicts phase transformations or microstructures from composition and processing variables. The second introduces intermediate physical variables, such as prior-austenite grain size, stored deformation energy, and thermodynamic quantities. The third automatically extracts microstructural representations from images and transfers them to property-prediction models. The first route is computationally efficient but relies strongly on data coverage when extrapolated. The second offers greater physical consistency but depends on the accuracy of auxiliary models and state variables. The third can capture spatial heterogeneity but remains constrained by labeling quality and imaging conditions. At present, the approach with the greatest engineering potential is not a single algorithm but an integrated framework combining physical state variables, image-based quantification, and industrial validation.

4.2. Mechanical Property Prediction

Machine-learning studies of structural-steel mechanical properties should be distinguished according to the target property because tensile strength, impact toughness, and elevated-temperature properties differ substantially in their physical definitions, testing procedures, data scatter, and validation requirements. The following discussion therefore separates room-temperature tensile properties, Charpy impact toughness, and elevated-temperature mechanical properties rather than comparing their prediction accuracy directly.
Room-temperature tensile properties. Yield strength, ultimate tensile strength, and elongation are the most frequently modeled room-temperature properties. Guo et al. analyzed 63,137 cleaned production records from an industrial steel-production database containing 27 composition- and process-related variables and developed simultaneous prediction models for yield strength, ultimate tensile strength, and elongation. The models were subsequently incorporated into a constrained nonlinear-programming framework to identify feasible regions for multiple property requirements [50]. However, the large record count should not be interpreted as an equivalent number of statistically independent material observations. The published study does not report the number of independent heats or production campaigns represented by the 63,137 records, and the validation was based on internal data partitioning rather than independent production-line validation. Accordingly, the reported performance should be interpreted primarily as evidence of interpolation within the represented production domain rather than as proof of cross-plant generalization. Target ranges and measurement uncertainty were also not standardized across the database, which limits direct comparison with independently generated datasets.
For high-dimensional industrial data with relatively few independent material states, introducing physically meaningful intermediate variables can be more effective than simply increasing model complexity. Jiang et al. transformed approximately 100 production parameters for SGLX82A pearlitic steel wire into microstructural descriptors, including proeutectoid ferrite fraction, pearlite fraction, and interlamellar spacing, using thermodynamic, kinetic, and finite-element calculations. The final dataset contained 150 instances after data cleaning. Gradient tree boosting and Gaussian process regression achieved mean relative errors below 0.7% and maximum relative errors below 2.0% under repeated 10-fold cross-validation, and the selected model was additionally evaluated using 10 unseen samples [51]. However, this result should not be interpreted as a universal industrial prediction accuracy because the dataset contained only 150 instances from a single production line and the reported unseen samples did not constitute demonstrated cross-production-line validation. Because the dataset originated from a single production line and a single steel-wire grade, these unseen samples provide evidence of within-domain prediction but should not be regarded as independent cross-production-line validation. The study also demonstrated that multiscale, domain-knowledge-based feature construction reduced prediction errors relative to purely data-based feature-selection strategies; nevertheless, a standardized naive baseline and experimental measurement uncertainty were not reported. The unusually small prediction error should therefore be interpreted in the context of the dataset size, production uniformity, target definition, and validation design rather than as a universal accuracy benchmark.
Physics-guided learning provides another strategy for reducing the dependence on large, labeled datasets. Cui et al. used a strengthening model to generate virtual yield-strength labels for otherwise unlabeled production data, pretrained a deep neural network using these virtual labels, and subsequently fine-tuned the model using experimentally measured data [52]. This approach illustrates the potential of combining mechanistic knowledge with machine learning, but the uncertainty introduced by the virtual labels should be distinguished from experimental measurement uncertainty. The resulting prediction performance is therefore more appropriately viewed as evidence for physics-guided data efficiency than as a direct comparison with purely experimental datasets.
Charpy impact energy represents a distinct prediction task because the measured response depends not only on composition and microstructure but also on specimen geometry, notch configuration, sampling orientation, and test temperature. Wu et al. compared shallow neural networks, extreme learning machines, and deep neural networks for predicting the Charpy V-notch impact energy of low-carbon steel. The Bayesian-optimized DNN achieved a correlation coefficient of 0.9536 and an RMSE of 17.34 J, with the modeled steel having a final thickness of 7.5 mm [53]. These values demonstrate the feasibility of machine-learning prediction within the investigated industrial dataset, but they should not be compared directly with errors from tensile-property models because the target variable, specimen configuration, and test conditions are fundamentally different. In particular, specimen dimensions, orientation, notch geometry, and test temperature should be treated as essential metadata when combining impact-test results from different sources.
Shang et al. further combined industrial production data with literature-derived data for pipeline steel and reported an improvement in the random-forest prediction performance from 0.58 to 0.90 when literature data were incorporated [54]. The result demonstrates the potential of literature-assisted learning to expand a limited industrial dataset, but it also highlights the risk of introducing systematic heterogeneity. Differences in Charpy specimen geometry, notch configuration, sampling orientation, test temperature, and property definitions can produce apparent model improvements that partly reflect changes in data distribution rather than genuine improvement in materials generalization. Therefore, impact-toughness models should be evaluated using harmonized testing metadata and, where possible, source- or specimen-level grouped validation.
Elevated-temperature mechanical properties. Elevated-temperature prediction constitutes a separate task because both the testing path and the definition of the target property influence the measured response. Shaheen et al. compiled 366 elevated-temperature test results from 19 experimental programs covering high-strength structural steels with nominal yield strengths from 460 to 960 MPa [55]. The database included both steady-state and transient-state tests, while heating rate and holding time were recorded when available. The study used chemical composition and temperature to predict the reduction factors for ultimate tensile strength, yield strength, 0.2% proof strength, and Young’s modulus. The reported data showed substantial scatter associated with testing method, manufacturing route, chemical composition, heating rate, holding time, and other incompletely reported experimental conditions. These differences should therefore be considered when interpreting machine-learning performance. In particular, elevated-temperature in-fire properties should not be conflated with post-fire residual properties, because the latter involve a different thermal history and require separate experimental validation.
Peng et al. incorporated CALPHAD-derived transformation temperatures, phase fractions, and precipitate fractions together with experimental composition, heat-treatment, temperature, and prior-austenite-grain-size variables to predict the elevated-temperature yield strength of 9–12Cr steels [56]. This study is included primarily as a methodological transfer case rather than direct evidence for conventional civil-infrastructure structural steels. Its main value lies in demonstrating how thermodynamic and microstructural descriptors can constrain machine-learning predictions and improve physical interpretability. However, the narrow steel-grade range and uncertainties associated with the auxiliary thermodynamic calculations limit direct extrapolation to other structural-steel families.
Across these three property classes, model performance should therefore be interpreted together with the statistical independence of samples, testing conditions, validation type, and uncertainty treatment. High accuracy obtained from randomly partitioned industrial records, repeated measurements from the same production campaign, or heterogeneous literature datasets do not constitute equivalent evidence of engineering generalization. For safety-critical structural-steel applications, external validation across independent heats, production campaigns, laboratories, or testing conditions should be regarded as stronger evidence than a larger internally randomized test set. Representative machine-learning studies for mechanical-property prediction are summarized in Table 5. As a methodological transfer case, Peng et al. incorporated thermodynamic and microstructural descriptors into the machine-learning feature space, as illustrated in Figure 5 [56].
The comparison indicates that reported prediction accuracy should be interpreted together with dataset independence, validation strategy, and testing conditions rather than used as a universal ranking of machine-learning algorithms.
Across these studies, no consistent algorithm ranking can be established because the datasets, target definitions, feature spaces, and validation protocols differ substantially. Tree-based models are frequently competitive for tabular datasets, but their apparent advantages remain dataset- and validation-dependent. The more defensible distinction is therefore between internally validated interpolation and externally demonstrated generalization. However, model rankings are strongly influenced by dataset size, feature construction, and validation strategy. Models trained on large industrial datasets may achieve high internal accuracy while learning production-line-specific correlations. Small-data models can reduce prediction errors by incorporating strengthening theories, CALPHAD-derived features, or pretraining strategies, but they simultaneously introduce uncertainty from the auxiliary physical models or virtual labels. High predictive accuracy therefore cannot replace independent validation across heats, manufacturers, production lines, and testing standards. Results lacking such external evidence should be interpreted as interpolation within a specific dataset rather than as universally applicable materials relationships.

4.3. Multi-Property Optimization and Inverse Materials Design

Structural-steel design is usually subject to simultaneous constraints on strength, ductility, toughness, weldability, durability, and cost. Maximizing a single property can therefore cause deterioration in other properties. Diao et al. established models for tensile strength, fracture strength, impact energy, hardness, fatigue strength, and elongation of carbon steels, and used the product of strength and elongation to characterize the strength–ductility balance. Combined with efficient global optimization, this approach illustrates the transition from single-property prediction toward multi-objective materials decision-making [57].
Forward property models combined with intelligent search algorithms have become a major route for composition-based inverse design. Lee et al. developed tensile-strength and total-elongation models using 1075 medium-Mn steel datasets and predicted that Fe–5.5Mn–0.2C–0.3Si steel austenitized at 780 °C could achieve a tensile strength of 1957 MPa and an elongation of 10.7%; the corresponding experimental values were 1952 MPa and 9.9%, respectively [58]. When the database was expanded to 1520 datasets containing microalloying elements and a genetic algorithm was introduced, millions of composition–austenitization combinations could be screened for candidates with tensile strengths above 2GPa. Five experimentally examined steels reached the target strength while retaining useful ductility [59]. These studies demonstrate the feasibility of combining predictive models with optimization, but agreement between predicted and measured properties alone does not establish engineering feasibility.
For engineering-oriented inverse design, candidate materials should additionally satisfy compositional tolerances, feasible processing windows, weldability, cost, and other manufacturing constraints. In particular, an optimized composition located in a sparse region of the training data may have a large prediction uncertainty and may be sensitive to small variations in composition or processing parameters. Therefore, Pareto-optimal candidates should be interpreted as a feasible set rather than a deterministic optimum. The objective functions, constraint boundaries, optimization domain, and treatment of predictive uncertainty should be explicitly considered when ranking candidates. This is especially important for structural steels, for which a small improvement in a target property may not justify reduced robustness or increased manufacturing cost.
Inverse-design targets have also expanded from alloy composition to microstructure. Wang and Adachi established models that predicted stress-strain curves, tensile strength, and total elongation from two- and three-dimensional microstructural descriptors and demonstrated the possibility of inferring microstructural parameters from target properties [60]. However, the inverse mapping from properties to microstructure is generally non-unique: different combinations of phase fraction, morphology, spatial distribution, orientation, and interface characteristics may produce similar macroscopic properties. Consequently, a candidate microstructure generated in a latent space should not be regarded as a directly realizable material unless it can be connected to a feasible composition and processing route.
Pei et al. used 621 SEM images of 9–12Cr ferritic-martensitic steels to construct a variational autoencoder-based representation of microstructure. The images were cropped into standardized sub-images, and the encoder, decoder, and regression network were used to extract latent microstructural features, reconstruct images, and predict corresponding alloying-element contents [61]. As illustrated in Figure 6, this approach demonstrates that microstructural images can serve not only as outputs for property interpretation but also as information carriers for inverse alloy design. Nevertheless, the reconstructed microstructure should be distinguished from a manufacturable microstructure, because the latter must also satisfy phase-transformation, processing, and compositional constraints.
Generative models and probabilistic optimization can further improve the exploration of complex microstructural spaces. Kusampudi and Diehl used a variational autoencoder to obtain low-dimensional representations of dual-phase steel microstructures and Bayesian optimization to generate structures with target yield strength and reduced damage sensitivity [62]. Lertkiatpeeti et al. combined representative-volume-element simulations, support-vector regression, neural networks, and Markov-chain Monte Carlo inversion, introducing spatial descriptors such as the Moran index, martensite banding index, and orientation to describe microstructural heterogeneity [63]. Their results demonstrate that identical phase fractions do not necessarily produce identical mechanical properties because phase connectivity, clustering, orientation, and interface distribution also influence the response. These findings further emphasize the non-uniqueness of inverse microstructure–property relationships and the need to incorporate physically realizable processing routes into inverse design.
Inverse-design objectives have also expanded to include cost and service performance. Allen et al. optimized alloy cost while satisfying Jominy hardenability constraints and achieved an average reduction in alloying cost of approximately 18% [64]. Studies of RAFM, austenitic stainless, and TWIP steels have further demonstrated multi-temperature strength–ductility optimization and simultaneous optimization of multiple mechanical properties [65,66]. Wang et al. incorporated strength, ductility, and corrosion resistance of medium-Mn steel within an interpretable optimization framework and experimentally evaluated the resulting candidates [67]. These studies indicate that the objective of multi-property design is not to identify a universally optimal composition, but to construct a set of candidates that simultaneously satisfy performance, physical, manufacturing, and service constraints.
An important distinction should also be made between retrospective and prospective validation. Recovering a previously reported alloy or reproducing a known property represents retrospective validation and mainly tests whether the model can identify solutions already contained in the available knowledge base. By contrast, prospective validation requires the optimization framework to select previously untested candidates, followed by independent material preparation, processing, microstructural characterization, and property testing. The latter provides stronger evidence for the practical value of inverse design because it evaluates the model in a genuinely predictive setting.
Machine-learning-assisted composition design generally consists of data preparation, property modeling, model interpretation, optimization, and experimental verification. Zhou et al. used 273 data entries for TWIP steels containing composition, microstructure, testing-condition, and mechanical-property information to develop prediction models for yield strength, ultimate tensile strength, and elongation. SHAP analysis was subsequently used to identify influential compositional and microstructural variables, followed by multi-objective optimization and experimental validation of newly designed TWIP steels [68]. As illustrated in Figure 7, this framework establishes a closed loop from historical data and property prediction to candidate selection and experimental verification.
Overall, inverse design in structural steels should be viewed as a constrained and uncertainty-aware decision process rather than a direct inversion of a forward prediction model. A reliable workflow should connect target properties with feasible composition, microstructure, and processing routes, quantify prediction uncertainty, and distinguish known-alloy recovery from prospective validation. Such an approach is more consistent with the requirements of structural-steel engineering, where robustness, manufacturability, weldability, cost, and service reliability are as important as the nominally optimized mechanical properties.

5. Applications of Machine Learning to Durability and Service-Life Assessment of Structural Steels

Machine learning applications in structural steels have increasingly extended from conventional property prediction to degradation assessment and service-life management. Unlike composition–property modeling, durability-related tasks involve evolving environmental conditions, time-dependent damage, multiple spatial scales, and different levels of engineering evidence. To avoid directly comparing fundamentally different prediction tasks, the studies reviewed in this section are distinguished according to degradation mechanism, target variable, temporal scale, and evidence level.
The principal tasks include corrosion-rate prediction, localized-damage assessment, elevated-temperature property degradation, member-level fire resistance, fatigue-strength and total-life prediction, crack-growth monitoring, and remaining-life estimation (Table 6). At the material level, models mainly predict corrosion rate, retained mechanical properties, or fatigue strength; at the component level, they address member resistance and crack evolution; and at the service level, they estimate damage progression, failure probability, or remaining life. Accordingly, reported prediction accuracy should be interpreted together with the physical target, observation period, validation strategy, and applicability domain rather than used to establish a universal ranking of machine-learning algorithms.
For corrosion and environmental degradation, particular attention is required to distinguish average or instantaneous corrosion from localized damage. Mass loss, corrosion current density, and average penetration rate describe overall degradation, whereas maximum pit depth and pit morphology are more directly related to local stress concentration and structural reliability. These quantities are therefore not interchangeable. Similarly, corrosion under high-temperature chemical environments is considered here as an environmental degradation process, whereas the thermomechanical response of structural steel during fire exposure is addressed separately in Section 5.2.

5.1. Corrosion and Environmental Degradation

Corrosion of structural steels is governed by the coupled effects of material composition, surface condition, environmental exposure, and time. Machine learning is particularly useful because it can simultaneously process material, environmental, and temporal variables and capture nonlinear interactions that are difficult to represent using conventional empirical models. However, corrosion prediction studies use substantially different target variables, including instantaneous corrosion current, average corrosion rate, corrosion current density, corrosion depth, and remaining service life. These quantities describe different stages or aspects of degradation and should therefore not be directly compared by their numerical prediction accuracy.
Early studies focused on short-term, continuously monitored atmospheric corrosion. Pei et al. monitored carbon steel using an Fe/Cu galvanic corrosion sensor for 34 days and used relative humidity, temperature, rainfall, airborne particles, and gaseous pollutants to predict instantaneous atmospheric corrosion. Random forest outperformed artificial neural networks and support vector regression, while explicitly considering rust formation further improved prediction. The study demonstrates the value of time-resolved sensor data but represents a relatively short exposure period and primarily describes the sensor-level corrosion response rather than long-term structural material loss [69].
Long-term atmospheric studies have incorporated both alloy composition and exposure-site information. Yan et al. used 306 corrosion records covering 18 alloy steels exposed at three marine-atmospheric sites for 1, 2, 3, 5, 7, and 10 years. The target was corrosion rate expressed as annual corrosion depth (μm·a−1). An optimized random forest model achieved R2 values of 0.94 and 0.73 for the training and testing sets, respectively [70]. These results demonstrate the importance of jointly considering material, environmental, and exposure-time variables, but the relatively small number of independent alloy instances means that the nominal record count should not be interpreted as equivalent to independent material observations.
The environments experienced by actual structures are rarely stationary. Song et al. investigated dynamic atmospheric corrosion under vehicle operating conditions by combining meteorological and pollution data with operational parameters, including average vehicle speed and the ratio of moving to stationary time. A time-weighting method was used to reconcile variables recorded at different sampling frequencies. Genetic-algorithm-optimized support vector regression achieved an R2 value of 0.9771, outperforming both the unoptimized support vector regression model and neural-network models [71]. Liu and Li predicted the corrosion rate of 3C steel in seawater using temperature, pH, dissolved oxygen, salinity, and oxidation–reduction potential as input variables. SHAP analysis identified oxidation–reduction potential as the dominant predictor [72]. These studies indicate that machine-learning-based corrosion analysis has progressed from fitting average environmental conditions toward modeling dynamic service scenarios and providing local explanations of individual predictions.
The application scope of corrosion prediction has also expanded from atmospheric and seawater environments to underground structures and steel–concrete systems. Dong et al. used 1428 experimental records to predict the corrosion current density of steel buried in soil. The inputs included exposure duration, moisture content, pH, electrical resistivity, chloride content, sulfate content, and total organic carbon. Random forest achieved an R2 value of 0.987, and soil resistivity, exposure duration, and total organic carbon formed the most effective three-variable combination [73].
Corrosion studies of reinforcing steel embedded in carbonated cementitious materials do not constitute primary evidence for structural-steel applications, but their treatment of multiphase material parameters, pore-solution chemistry, and environmental humidity provides transferable methodological insights. Ji and Ye identified electrical resistivity, the [Cl]/[OH] ratio, cement proportion, and corrosion potential as critical variables [74]. However, the resulting model cannot be directly extrapolated to exposed structural steel or reinforcing steel in non-carbonated concrete. This limitation demonstrates that corrosion models must explicitly define the substrate, surface condition, and environmental boundaries within which they are valid. Figure 8 summarizes the general workflow from structural-steel corrosion data to service-life assessment.
Figure 8 presents a sequential framework comprising material information, environmental and operational conditions, corrosion monitoring, corrosion-state prediction, and structural service-life assessment. Material information includes chemical composition, microstructure, surface condition, and protective coatings. Environmental variables include temperature, humidity, chloride deposition, pH, pollutants, and wet-dry cycling. Monitoring data may include mass loss, corrosion current, electrical resistance, wall thickness, and corrosion morphology. The model outputs progress from instantaneous corrosion rate and cumulative corrosion depth to cross-sectional loss, load-carrying-capacity degradation, and remaining service life.
Under high-temperature molten-salt, oxidation, and thermochemical environments, corrosion is additionally coupled with temperature, alloy composition, corrosive medium, and exposure duration. Muthukrishnan et al. compiled high-temperature corrosion data for structural materials and protective coatings and compared multiple regression methods. Random forest generally exhibited favorable performance. However, the study also emphasized that currently available data do not adequately cover different molten-salt chemistries, temperature ranges, and material states. The resulting models are therefore more appropriate for preliminary material screening than for replacing long-term corrosion testing [75].
The engineering objective of corrosion prediction is to translate environmental conditions and corrosion states into estimates of cross-sectional loss, structural reliability, and maintenance timing. Ensemble models based on pipeline-accident records and a physics-informed neural network (PINN) developed for marine steel pipe piles represent two different routes for remaining-life prediction [76,77]. The former depends strongly on the representativeness of historical accident databases, whereas the latter relies on assumptions concerning uniform corrosion and boundary conditions. Neither approach yet fully captures the combined effects of extreme pitting, scour, and corrosion fatigue.
Feng et al. incorporated chloride diffusion, corrosion depth, wall-thickness loss, and mechanical-property degradation into a PINN for service-life prediction of marine steel pipe piles. After particle swarm optimization, the deviation between the inversely identified diffusion coefficient and the experimental value was 4.3% [77]. Nevertheless, the model approximated the chloride penetration front as a uniform corrosion front. Its predictions therefore primarily represent average wall-thickness degradation and do not fully capture pitting corrosion, crevice corrosion, or corrosion-scour interactions. Figure 9 and Table 7 present, respectively, the measured and machine-learning-predicted long-term corrosion rates of low-alloy steels and additional representative studies on corrosion and environmental degradation of structural steels.
Across exposure classes, the prediction targets and validation strategies differ substantially, making direct comparisons of reported accuracy inappropriate. Short-term atmospheric studies mainly address instantaneous electrochemical responses, whereas long-term atmospheric and soil studies focus on average corrosion rates. Service-life models further transform corrosion degradation into wall-thickness loss or failure time. These tasks should therefore be evaluated according to their physical target, temporal scale, validation design, and engineering consequence rather than by a single accuracy metric.

5.2. Fire and Extreme-Temperature Performance

Fire and extreme-temperature performance involves a hierarchical transfer from material degradation to member resistance and finally to structural or infrastructure-level consequences. Temperature-dependent thermal and mechanical properties determine the constitutive response of steel; these properties, together with geometry, loading, thermal gradients, and boundary conditions, determine member stability and fire resistance; member failures can subsequently affect structural robustness and infrastructure-level fire risk. Member- and infrastructure-scale studies are therefore included because they demonstrate how material degradation is translated into engineering consequences. However, they are not treated as direct evidence of material behavior, and their datasets and performance metrics are not directly compared with material-level models.
Predicting temperature-dependent material properties represents only the first level of machine-learning-based fire assessment. Naser used ANN and genetic algorithms to integrate experimental data and material models from fire codes and standards, including ASCE, Eurocode and British Standards, to derive temperature-dependent thermal and mechanical properties of structural steel [78]. The predicted quantities included thermal conductivity, specific heat, yield strength, and elastic modulus or their reduction factors. The study demonstrates that machine learning can reconcile differences among heterogeneous material models, but the resulting expressions remain dependent on the definitions and testing assumptions contained in the source data.
The transferability of temperature-dependent material models is strongly affected by the experimental protocol. Steady-state and transient-state tests do not necessarily produce identical stress–strain responses, while heating rate, strain rate, thermal exposure duration, cooling regime, and post-fire testing conditions can influence the measured properties [78,79]. Accordingly, room-temperature-to-elevated-temperature prediction, transient fire response, and post-fire residual-property prediction should be regarded as distinct tasks rather than a single temperature–property mapping.
The next evidence level concerns structural members, where temperature-dependent material properties interact with geometry, slenderness, loading, eccentricity, boundary conditions, and instability modes. Zhao developed a hybrid neural network optimized by a genetic algorithm to predict the failure temperature of steel columns, incorporating geometric and loading parameters [80]. The study demonstrates that machine learning can capture the combined influence of material and member parameters, but comparison with analytical design equations remains essential because a lower statistical error does not automatically imply a safer engineering prediction.
Possidente and Couto used 21,879 geometrically and materially nonlinear finite-element simulations to train neural-network, random-forest, and support-vector-machine models for compressed L-, T-, and X-shaped steel members at elevated temperatures [81]. The large numerical dataset enables efficient surrogate prediction, but the samples are simulation-derived rather than independent fire-test observations. The models therefore provide a useful numerical proxy while retaining the assumptions and model-form uncertainty of the underlying finite-element simulations. The study itself emphasizes the trade-off between accuracy and safety and shows that applicability is limited by the section shapes and heating conditions represented in the training data.
These member-level models demonstrate that algorithm selection for safety-critical applications cannot be based solely on average error metrics. A model may achieve a low overall root mean square error while still producing non-conservative predictions in regimes involving high temperatures, large slenderness ratios, or torsional buckling. Machine-learning models for member fire resistance should therefore report prediction bias within critical regions, the proportion of unconservative estimates, and quantile-based errors, while also considering the safety margins embedded in design codes. Figure 10 summarizes the multiscale prediction process from material degradation to structural fire risk.
Figure 10 sequentially links fire scenarios and temperature histories, degradation of thermal and mechanical properties, member stability and resistance, structural failure modes, and fire-risk classification. Data at the different levels are derived from elevated-temperature material tests, finite element simulations, member fire tests, and historical fire-incident records.
Machine learning has also been applied to global failure and progressive-collapse assessment of steel structures. Fu used Monte Carlo simulation and random sampling to generate data on fire scenarios, member temperatures, load ratios, and critical temperatures. Decision trees, k-nearest neighbors, and artificial neural networks were then compared for classifying the failure modes of steel-frame members. The predicted member failures were subsequently used to remove failed components from the structural model and assess the likelihood of progressive collapse [82].
For infrastructure-level risk screening, Kodur and Naser analyzed fire incidents involving 80 steel bridges and 38 concrete bridges using random forest, support vector machines, and generalized additive models (GAMs). Input variables included bridge material, structural form, span length, age, number of traffic lanes, geographical importance, fuel type, and fire location. The overall classification accuracy was approximately 70%, and the different models consistently identified fuel type, span length, or bridge age as influential factors [83]. Figure 11 presents the machine-learning-based bridge fire-risk assessment workflow reported in the open-access study.
Another major limitation of machine learning for fire engineering is the scarcity of high-quality data. Naser et al. developed the StructuresNet and FireNet benchmark datasets and used them to compare decision trees, random forest, extreme gradient boosting, Light Gradient Boosting Machine (LightGBM), and deep-learning algorithms under consistent data and evaluation protocols. Their work emphasized that standardized datasets and unified evaluation procedures are prerequisites for meaningful comparisons among fire-performance models [84].
Table 8 summarizes representative machine-learning applications for elevated-temperature and fire-performance assessment of structural steels. Generative models may alleviate the shortage of full-scale fire-test data. However, the available representative study focused on reinforced-concrete columns and should therefore be regarded only as a methodological transfer case for data augmentation [85]. Synthetic data can increase sample density within the existing training distribution, but they cannot automatically generate buckling, connection failure, or progressive-collapse mechanisms that are absent from real experiments or reliable numerical simulations. Generative adversarial network (GAN)- and VAE-based augmentation must therefore be validated jointly against physics-based models and independent fire tests.
Current machine-learning studies on structural fire performance remain characterized by substantial disconnection across scales. Material-level models commonly neglect member boundary conditions and non-uniform temperature fields, member-level models rely heavily on numerical simulations, and structural-risk models often lack information on steel grade and fire-protection systems. A more rational direction is to connect temperature fields, material constitutive behavior, member failure, and structural-system response within a hierarchical framework and to propagate predictive uncertainty across each level, rather than developing isolated black-box models for individual stages.

5.3. Fatigue, Fracture, and Remaining-Life Assessment

Fatigue and fracture assessment of structural steels covers several distinct tasks, including fatigue-strength prediction, S-N curve construction, total fatigue-life prediction, low-cycle fatigue, welded-joint fatigue, corrosion fatigue, crack-growth-rate prediction, crack detection, and remaining-life estimation. These tasks differ in their input variables, target definitions, temporal scales, and evaluation criteria and therefore should not be compared solely by a common prediction error. In particular, fatigue-strength and S-N predictions generally focus on stress–life relationships, whereas crack-growth models predict da/dN, vision or nondestructive evaluation (NDE) models identify or quantify damage, and remaining-life models estimate the future failure horizon. This distinction is essential when assessing the engineering relevance of machine-learning (ML) models.
Early studies based on NIMS structural-steel fatigue data demonstrated that chemical composition and manufacturing parameters could be used to predict fatigue strength and identify influential variables [86,87]. Arvanitis et al. subsequently used NIMS data together with supplementary experiments to compare polynomial regression, support vector regression, XGBoost, and neural networks for structural-steel fatigue-life prediction. XGBoost achieved the lowest mean squared error, while polynomial regression remained competitive with neural networks at substantially lower computational cost [88]. This result indicates that, for structured fatigue datasets of moderate size, increasing model complexity does not necessarily lead to better engineering prediction. Figure 12 presents the general workflow for machine-learning-based fatigue-life prediction.
A critical issue in fatigue databases is the treatment of run-outs and right-censored observations. Specimens that have survived a prescribed number of cycles without failure do not provide an exact failure life and should not simply be discarded or treated as failed at the test limit. Their exclusion may bias the predicted high-cycle fatigue region, whereas assigning the test limit as an exact failure can distort the upper tail of the life distribution. Therefore, fatigue ML studies should explicitly report the number of run-outs, their censoring treatment, and, where appropriate, use censored regression or survival-analysis approaches. This consideration is particularly important when comparing models for total fatigue life and remaining-life prediction.
Welded joints require additional consideration because local geometry and manufacturing history strongly influence fatigue behavior. Relevant variables include joint class, weld geometry, weld-toe radius, defect type and size, residual stress, post-weld treatment, plate thickness, loading mode, stress ratio, mean stress, and variable-amplitude loading spectra. Feng et al. first transformed physical parameters from a welded-fatigue database and used XGBoost to determine feature weights before incorporating these weights into a deep convolutional neural network for S-N prediction [89]. Schubnell et al. further applied transfer learning from non-welded steel specimens to welded joints and compared the resulting predictions with conventional ML and fracture-mechanics-based approaches [90]. These studies indicate that transfer learning can reduce the data requirement of welded-joint prediction, but its effectiveness depends on whether the source and target domains share physically meaningful fatigue variables.
For engineering assessment, ML models should therefore be compared not only with alternative algorithms but also with established physical or design baselines. Depending on the task, appropriate baselines include S-N design curves, local stress/strain approaches, empirical fatigue relationships, Paris-law-type crack-growth models, and fracture-mechanics-based life assessment. Agreement with an ML model alone does not establish engineering validity; the additional value of ML should be demonstrated through improved prediction, broader applicability, reduced experimental requirements, or more efficient updating under new loading or environmental conditions.
Environmental effects further increase the variability of fatigue behavior. Feng et al. combined material strength, loading frequency, stress ratio, loading mode, temperature, and surface condition with Borderline-SMOTE, XGBoost, and an attention-based DCNN to predict corrosion-fatigue behavior [91]. The reported results show that incorporating environmental and mechanical variables can improve prediction under different test conditions. However, corrosion-fatigue models should distinguish between average degradation and localized damage because pits, corrosion defects, and crack-initiation sites may control structural failure even when the average corrosion rate remains relatively low. Validation should therefore preserve the chronological or site-specific structure of environmental data rather than relying solely on random record-level splitting.
Low-cycle fatigue provides another methodological transfer case. Jiang et al. incorporated temperature and strain rate into a physics-informed neural network for 316 stainless steels and introduced physical constraints into the loss function [92]. The predicted fatigue lives were reported within the twice-error band. Although this approach demonstrates the value of physical constraints for small datasets, the material-specific cyclic softening, creep-fatigue interaction, and temperature dependence of 316 stainless steel should not be assumed to be directly transferable to structural steels. The main methodological implication is that physical constraints can regularize ML models, but the constraints themselves must be reconstructed for the target steel and loading regime.
Fatigue-crack-growth prediction provides a direct connection between damage evolution and remaining-life assessment. Kamble et al. used carbon-steel compact-tension data to compare regression and nearest-neighbor approaches for predicting crack-growth behavior across both stable and rapid-growth regimes [93]. Such models extend beyond the conventional use of Paris-law-type relationships in the stable crack-growth region, but their engineering application still requires careful definition of the crack-size range, loading conditions, and fracture-mechanics variables. In particular, extrapolation beyond the experimentally covered crack-growth regime should not be interpreted as validated remaining-life prediction.
The development of image-based and NDE methods has further expanded ML from offline life prediction toward online damage detection. Long et al. used a global–local Faster R-CNN framework with images acquired during cyclic loading to identify small cracks and estimate crack length and growth rate [94]. Zhang et al. combined mechanoluminescent materials, computer vision, and ML for real-time fatigue-crack monitoring [95]. These studies demonstrate the potential of ML for surface-crack detection and quantitative monitoring. However, crack detection and remaining-life prediction are not equivalent tasks. Optical and other vision-based methods primarily provide information on visible surface damage and may not capture internal crack morphology. Future structural-steel applications should therefore combine multiple NDE modalities, such as ultrasonic testing, acoustic emission, magnetic methods, and X-ray computed tomography, where appropriate. Multimodal sensing may improve damage characterization, but the resulting measurements must still be linked to a physically meaningful damage state before they are used for prognostic life estimation [96].
A further methodological transfer case concerns hydrogen-assisted fatigue crack growth. For Cr-Mo steel exposed to gaseous hydrogen, gradient boosting combined with SHAP was used to evaluate the effects of stress-intensity-factor range, hydrogen pressure, strength, stress ratio, frequency, and chemical composition on crack-growth rate [97]. The study illustrates how ML can jointly represent material, loading, and environmental variables. However, the reported feature contributions should be interpreted only within the investigated Cr-Mo steel, hydrogen-pressure range, and loading conditions and should not be generalized as universal rankings for structural steels.
Another important issue is sample and sequence leakage. Fatigue datasets often contain repeated measurements from the same specimen, multiple images obtained during one test, or successive crack-growth observations from a single loading history. If these related observations are randomly divided between training and test sets, the model may partially recognize the specimen or test sequence rather than learn a transferable fatigue relationship. Therefore, specimen-level, test-level, or experiment-level grouping should be maintained during data partitioning, with all repeated observations from one specimen or test retained within the same partition. For crack-growth and image-based studies, this requirement is particularly important because adjacent measurements are strongly correlated. Figure 13 summarizes the continuous application chain of machine learning across fatigue-life prediction, crack-growth modeling, damage monitoring, and remaining-life assessment. Additional representative studies in this area are listed in Table 9.
Overall, ML applications in structural-steel fatigue have progressed from fatigue-strength and total-life prediction toward welded-joint assessment, environmental fatigue, crack-growth modeling, damage detection, and remaining-life estimation. However, these tasks require different target variables, baselines, validation strategies, and uncertainty treatments. The principal challenge is therefore not simply improving prediction accuracy, but establishing whether an ML model can provide a reliable engineering estimate under unseen specimens, loading histories, environments, and damage states. Future research should integrate fatigue databases, fracture-mechanics variables, multimodal NDE observations, and time-dependent monitoring data into dynamically updated remaining-life models.

6. Current Challenges and Future Directions

The studies reviewed in this study demonstrate that machine learning has extended from composition–property prediction to process optimization, microstructure characterization, inverse materials design, and durability and service-life assessment of structural steels. However, the principal barrier to further development is no longer whether high predictive accuracy can be achieved on an individual dataset, but whether the resulting models remain reliable when transferred across steel grades, production routes, laboratories, and service environments. For safety-critical structural steels, engineering deployment requires evidence of data independence, physical consistency, calibrated uncertainty, and performance beyond the original training domain.

6.1. Current Challenges

The first unresolved issue is the quality and independence of available evidence. Many structural-steel datasets contain large numbers of records but substantially fewer independent heats, specimens, production campaigns, exposure sites, or fatigue tests. As demonstrated in Section 3 and Section 4, internally randomized validation remains common, while independent cross-manufacturer, cross-laboratory, or field validation is much less frequent. Consequently, reported model accuracy often represents interpolation within an existing production or experimental domain rather than demonstrated engineering generalization. Data scarcity is particularly severe for fracture, localized corrosion, fire, corrosion fatigue, hydrogen-assisted cracking, and long-term monitoring, where the most safety-relevant observations are also the most difficult to obtain [97,98]. Inconsistent metadata, interoperability, and reporting practices further restrict reliable reuse and cross-source integration of materials data [99,100,101].
A second challenge is the limited connection between statistical prediction and physical state evolution. Composition and process variables alone may be sufficient for interpolation within a narrow steel family, but they do not uniquely describe microstructure, defects, residual stress, environmental damage, or crack evolution. Conversely, image-based and sensor-based models can capture local information but remain sensitive to specimen preparation, acquisition conditions, labeling, and spatial or temporal correlation. The major unresolved problem is therefore not simply multimodal data fusion, but the construction of scale-bridging state variables that connect composition, processing, microstructure, defects, local damage, and structural response without losing physical interpretability [102].
The third challenge concerns uncertainty and engineering conservatism. Most reported models still provide deterministic predictions and average performance metrics, whereas structural-steel applications require knowledge of whether an individual prediction is sufficiently reliable for a safety-related decision. Aleatoric variability in material properties and epistemic uncertainty caused by sparse training data must be distinguished. Moreover, prediction errors are direction-sensitive: overestimating fatigue life or member resistance can be more critical than an error of the same magnitude in the conservative direction. Applicability-domain assessment, calibrated prediction intervals, out-of-distribution detection, and explicit reporting of non-conservative errors therefore remain essential gaps between academic model accuracy and engineering reliability [103].
Finally, verification, governance, and deployment procedures remain immature. Many studies terminate after internal model testing, whereas few demonstrate prospective alloy production, independent manufacturing-line validation, full-scale structural testing, or long-term field deployment. Model updates caused by new steel grades, process drift, sensor replacement, or changing environmental conditions are rarely governed by predefined procedures. Version control, traceable audit records, responsibility for model updates, conservative fallback rules, cybersecurity of industrial data, and compatibility with existing material specifications and structural design standards must therefore be treated as part of model reliability rather than as downstream software issues.

6.2. Future Research Directions

In the near term, priority should be given to standardized data and reporting practices and to stronger validation protocols. Structural-steel datasets should preserve heat number, product form, thickness, sampling position and orientation, manufacturing and welding history, specimen geometry, testing standard, environmental conditions, censoring information, measurement uncertainty, and data-processing lineage. Reporting should distinguish raw records from statistically independent material units. Validation should progress from random splitting toward heat-, specimen-, source-, site-, and production-based grouping, followed where possible by independent manufacturer or laboratory data. Success should be assessed not only through lower RMSE or higher R2, but through predefined criteria such as successful external validation, calibrated prediction-interval coverage, out-of-distribution detection performance, and the frequency of non-conservative errors.
In the medium term, research should focus on physics-informed, multi-fidelity, transfer-learning, and uncertainty-aware frameworks. CALPHAD, phase-transformation kinetics, strengthening models, finite-element analysis, corrosion transport, and fracture mechanics can provide intermediate state variables, physical constraints, or lower-fidelity information, while machine learning can represent relationships that remain difficult to formulate explicitly [104,105,106]. Transfer learning should be accepted only when the source and target domains share identifiable physical information, and simulation-derived data should be calibrated against experimental observations rather than pooled indiscriminately with them. Active learning can further reduce experimental demand by selecting experiments with high expected information gain [107,108]. For materials design, a meaningful success criterion is not merely rediscovery of known alloys but prospective validation of genuinely new candidates, together with evidence of manufacturing robustness and reduced experimental campaigns relative to conventional search strategies.
Closed-loop experimentation represents a further medium-term objective. Systems such as CAMEO demonstrate how active learning, physical knowledge, and rapid characterization can iteratively update material-selection decisions [109]. For structural steels, closed-loop development should initially focus on laboratory alloy design, heat-treatment windows, welding parameters, and high-cost corrosion or fatigue experiments. Model-generated candidates should remain subject to expert constraints on manufacturability, weldability, cost, standards compliance, and safety. A closed loop should therefore be evaluated by the number of experiments required to reach a validated design, its ability to identify candidates outside the initial dataset, and the reproducibility of the resulting material rather than solely by optimization speed.
In the long term, digital twins could connect steel manufacturing history with evolving in-service damage and maintenance decisions. In this context, a structural-steel digital twin should be defined operationally rather than as a generic digital representation. Relevant state variables may include microstructure, local mechanical properties, residual stress, corrosion depth, pit geometry, crack size, and accumulated fatigue damage. Sensor inputs may include temperature, strain, corrosion measurements, ultrasonic or acoustic-emission signals, and crack-monitoring data. The twin should specify when state estimates are updated, how model parameters are recalibrated after new inspections, how uncertainty is propagated from sensing to remaining-life prediction, and which engineering decision is supported—for example, continued service, inspection scheduling, repair, or replacement. A useful digital twin must therefore demonstrate that dynamic updating improves decision reliability relative to a static design-stage life estimate [110]. Figure 14 presents a roadmap for trustworthy machine-learning deployment across structural-steel design, manufacturing, and in-service management.
Long-term deployment will also require explicit model governance and integration with engineering standards. Each deployed model should have a controlled version, traceable training and validation data, an audit trail of changes, assigned responsibility for approval and updating, predefined out-of-domain and sensor-failure responses, and a conservative fallback to an accepted physical model or design provision when confidence is insufficient. Cybersecurity and access control are particularly relevant when production records and field-monitoring data are incorporated into continuously updated models.
The National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework should therefore be regarded as a general governance reference rather than a structural-steel design standard [111]. Its principles can be translated into materials-specific verification activities: governance corresponds to model ownership, version control, and auditability; mapping corresponds to defining the steel grade, process, service condition, and consequence of failure; measurement corresponds to independent validation, calibration, applicability-domain assessment, and non-conservative error analysis; and management corresponds to model updating, out-of-distribution handling, conservative fallback, and retirement. Regulatory adoption will require these procedures to be connected to existing material qualification and structural-design practices rather than replacing established engineering standards with a generic ML model. The specific research gaps that need to be addressed for each of these areas are listed in Table 10.
Overall, the development pathway is expected to progress from standardized and independently validated datasets in the near term, through physics-informed and uncertainty-aware models with prospective experimental validation in the medium term, toward closed-loop material development, operational digital twins, and code-compatible model governance in the long term. The central measure of progress should therefore shift from dataset-specific prediction accuracy to demonstrated reductions in uncertainty, experimental burden, non-conservative prediction, and life-cycle decision risk.

7. Conclusions

This review shows that machine learning has become a useful complementary tool for structural-steel research across composition and process modeling, microstructure characterization, mechanical-property prediction, inverse design, and durability assessment. The available evidence indicates that tree-based methods are frequently effective for small- to medium-sized tabular datasets, whereas image- and sequence-based tasks increasingly employ deep-learning approaches. However, model performance is strongly dependent on dataset composition, feature representation, and validation strategy; no algorithm can therefore be regarded as universally superior across structural-steel applications.
The maturity of the available evidence differs substantially among application areas. Composition, processing, and mechanical-property prediction are comparatively well developed, and a limited number of alloy- and process-design studies have progressed to experimental validation of newly selected candidates. Some industrial-property models have also been evaluated using previously unseen production samples. In contrast, corrosion, fire, fatigue, crack-growth, and remaining-life models are still dominated by internal validation, literature-derived datasets, laboratory-scale experiments, or simulation-based evidence. Prospective cross-manufacturer validation, independent production-line testing, full-scale structural verification, and long-term field deployment remain uncommon. Consequently, many service-life applications should presently be regarded as proof-of-concept or decision-support tools rather than validated replacements for established engineering models or design provisions.
Across the reviewed studies, the most persistent limitation is not predictive accuracy itself but the strength of the supporting evidence. Large record counts often do not correspond to equally large numbers of independent heats, specimens, joints, or exposure campaigns, and random data partitioning can obscure this distinction. For safety-critical use, future progress should therefore be judged by independent validation, calibrated uncertainty, applicability-domain control, physically meaningful baselines, and the rate of non-conservative predictions rather than by R2, RMSE, or classification accuracy alone. Physics-informed descriptors, multi-fidelity learning, and multimodal monitoring are promising only when their added complexity produces demonstrable improvements in external generalization or engineering decision quality.
This review also has limitations. The synthesis is affected by publication and literature-selection bias, heterogeneity in prediction targets and performance metrics, and incomplete reporting of independent sample numbers, testing conditions, uncertainty, and external validation in many primary studies. In addition, selected adjacent steel systems and methodological transfer cases were included to illustrate approaches not yet sufficiently demonstrated in conventional civil-infrastructure structural steels; these examples should not be interpreted as direct evidence of structural-steel readiness. Overall, the field has progressed beyond proof-of-principle property fitting, but widespread engineering deployment will require a shift from dataset-specific accuracy toward independently validated, uncertainty-aware, and physically defensible models.

Author Contributions

Methodology, G.W. and M.L.; Data curation, W.X. and A.M.S.; Writing—original draft preparation, G.W. and M.L.; Writing—review and editing, W.X. and A.M.S.; Supervision, B.C.; Resources, B.C. All authors have read and agreed to the published version of the manuscript.

Funding

Interdisciplinary Research Project for Disciplinary Cluster Development of Jilin Agricultural Science and Technology College, (2026)JX403, Guomin Wei.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Hkdh, B. Neural networks in materials science. ISIJ Int. 1999, 39, 966–979. [Google Scholar] [CrossRef] [Scilit]
  2. Müller, M.; Stiefel, M.; Bachmann, B.I.; Britz, D.; Mücklich, F. Overview: Machine Learning for Segmentation and Classification of Complex Steel Microstructures. Metals 2024, 14, 553. [Google Scholar] [CrossRef] [Scilit]
  3. Miao, B.; Lin, G.Q.; Zhang, Y.; Cheng, Y.W.; Yang, H. Machine learning applications in metallic materials: Recent advances and future perspectives. J. Alloys Compd. 2026, 1062, 187486. [Google Scholar] [CrossRef] [Scilit]
  4. Hart, G.L.; Mueller, T.; Toher, C.; Curtarolo, S. Machine learning for alloys. Nat. Rev. Mater. 2021, 6, 730–755. [Google Scholar] [CrossRef] [Scilit]
  5. Hu, M.; Tan, Q.; Knibbe, R.; Xu, M.; Jiang, B.; Wang, S.; Li, X.; Zhang, M.X. Recent applications of machine learning in alloy design: A review. Mater. Sci. Eng. R Rep. 2023, 155, 100746. [Google Scholar] [CrossRef] [Scilit]
  6. Pourrahimi, S.; Hakimian, S. Machine Learning for Alloy Design: A Property-Oriented Review. Alloys 2026, 5, 7. [Google Scholar] [CrossRef] [Scilit]
  7. Butler, K.T.; Davies, D.W.; Cartwright, H.; Isayev, O.; Walsh, A. Machine learning for molecular and materials science. Nature 2018, 559, 547–555. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Ramprasad, R.; Batra, R.; Pilania, G.; Mannodi-Kanakkithodi, A.; Kim, C. Machine learning in materials informatics: Recent applications and prospects. npj Comput. Mater. 2017, 3, 54. [Google Scholar] [CrossRef] [Scilit]
  9. Schmidt, J.; Marques, M.R.; Botti, S.; Marques, M.A. Recent advances and applications of machine learning in solid-state materials science. npj Comput. Mater. 2019, 5, 83. [Google Scholar] [CrossRef] [Scilit]
  10. Jha, D.; Ward, L.; Paul, A.; Liao, W.K.; Choudhary, A.; Wolverton, C.; Agrawal, A. Elemnet: Deep learning the chemistry of materials from only elemental composition. Sci. Rep. 2018, 8, 17593. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Xie, T.; Grossman, J.C. Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties. Phys. Rev. Lett. 2018, 120, 145301. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Choudhary, K.; DeCost, B.; Chen, C.; Jain, A.; Tavazza, F.; Cohn, R.; Park, C.W.; Choudhary, A.; Agrawal, A.; Billinge, S.J.L.; et al. Recent advances and applications of deep learning methods in materials science. npj Comput. Mater. 2022, 8, 59. [Google Scholar] [CrossRef] [Scilit]
  13. Lian, G.; Liu, X.; Wang, Q.; Shen, C.; Wang, Y.; Mu, W. Artificial intelligence-assisted non-metallic inclusion particle analysis in advanced steels using machine learning: A review. Int. J. Miner. Metall. Mater. 2026, 33, 401–416. [Google Scholar] [CrossRef] [Scilit]
  14. Voulgaris, S.; Chandrinos, S.; Chamatidis, I.; Kazakis, G.; Tsakalis, P.; Mitropoulou, C.C.; Georgantzinos, S.K.; Istrati, D.; Lagaros, N.D. Machine learning and data-driven methods for steel constitutive modeling: A state-of-the-art review and validation. Discov. Civ. Eng. 2025, 2, 228. [Google Scholar] [CrossRef] [Scilit]
  15. Wang, H.; Li, B.; Gong, J.; Xuan, F.Z. Machine learning-based fatigue life prediction of metal materials: Perspectives of physics-informed and data-driven hybrid methods. Eng. Fract. Mech. 2023, 284, 109242. [Google Scholar] [CrossRef] [Scilit]
  16. Hamada, A.; Elyamny, S.; Abd-Elaziem, W.; Elkatatny, S.; Darwish, M.A.; Sebaey, T.A.; Järvenpää, A.; Vineesh, K.; Elsheikh, A.H. Advancing fatigue life prediction with machine learning: A review. Mater. Today Commun. 2025, 43, 111525. [Google Scholar] [CrossRef] [Scilit]
  17. Stergiou, K.; Ntakolia, C.; Varytis, P.; Koumoulos, E.; Karlsson, P.; Moustakidis, S. Enhancing property prediction and process optimization in building materials through machine learning: A review. Comput. Mater. Sci. 2023, 220, 112031. [Google Scholar] [CrossRef] [Scilit]
  18. Sarfarazi, S.; Mascolo, I.; Modano, M.; Guarracino, F. Application of artificial intelligence to support design and analysis of steel structures. Metals 2025, 15, 408. [Google Scholar] [CrossRef] [Scilit]
  19. Fang, W.; Huang, J.X.; Peng, T.X.; Long, Y.; Yin, F.X. Machine learning-based performance predictions for steels considering manufacturing process parameters: A review. J. Iron Steel Res. Int. 2024, 31, 1555–1581. [Google Scholar] [CrossRef] [Scilit]
  20. Geng, X.; Wang, F.; Wu, H.H.; Wang, S.; Wu, G.; Gao, J.; Zhao, H.; Zhang, C.; Mao, X. Data-driven and artificial intelligence accelerated steel material research and intelligent manufacturing technology. Mater. Genome Eng. Adv. 2023, 1, e10. [Google Scholar] [CrossRef] [Scilit]
  21. Zhang, R.; Yang, J. State of the art in applications of machine learning in steelmaking process modeling. Int. J. Miner. Metall. Mater. 2023, 30, 2055–2075. [Google Scholar] [CrossRef] [Scilit]
  22. Tsutsui, K.; Namba, T.; Kihara, K.; Hirata, J.; Matsuo, S.; Ito, K. Current trends on deep learning techniques applied in iron and steel making field: A review. ISIJ Int. 2024, 64, 1619–1640. [Google Scholar] [CrossRef] [Scilit]
  23. Iren, D.; Ackermann, M.; Gorfer, J.; Pujar, G.; Wesselmecking, S.; Krupp, U.; Bromuri, S. Aachen-Heerlen annotated steel microstructure dataset. Sci. Data 2021, 8, 140. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. DeCost, B.L.; Hecht, M.D.; Francis, T.; Webler, B.A.; Picard, Y.N.; Holm, E.A. UHCSDB: Ultrahigh carbon steel micrograph database: Tools for exploring large heterogeneous microstructure datasets. Integr. Mater. Manuf. Innov. 2017, 6, 197–205. [Google Scholar]
  25. Lee, G.; Yi, G.H.; Tiong, L.C.O.; Sohn, S.S.; Kim, D. Internal defect database of mechanically deformed ferritic steel via X-ray computed tomography. Sci. Data 2025, 12, 1871. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. DeCost, B.L.; Holm, E.A. A computer vision approach for automated analysis and classification of microstructural image data. Comput. Mater. Sci. 2015, 110, 126–133. [Google Scholar] [CrossRef] [Scilit]
  27. Bostanabad, R.; Zhang, Y.; Li, X.; Kearney, T.; Brinson, L.C.; Apley, D.W.; Liu, W.K.; Chen, W. Computational microstructure characterization and reconstruction: Review of the state-of-the-art techniques. Prog. Mater. Sci. 2018, 95, 1–41. [Google Scholar] [CrossRef] [Scilit]
  28. Jha, D.; Choudhary, K.; Tavazza, F.; Liao, W.K.; Choudhary, A.; Campbell, C.; Agrawal, A. Enhancing materials property prediction by leveraging computational and experimental data using deep transfer learning. Nat. Commun. 2019, 10, 5316. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data 2016, 3, 160018. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Soedarmadji, E.; Stein, H.S.; Suram, S.K.; Guevarra, D.; Gregoire, J.M. Tracking materials science data lineage to manage millions of materials experiments and analyses. npj Comput. Mater. 2019, 5, 79. [Google Scholar] [CrossRef] [Scilit]
  31. Ma, Y.; Xu, P.; Li, M.; Ji, X.; Zhao, W.; Lu, W. The mastery of details in the workflow of materials machine learning. npj Comput. Mater. 2024, 10, 141. [Google Scholar] [CrossRef] [Scilit]
  32. Ward, L.; Agrawal, A.; Choudhary, A.; Wolverton, C. A general-purpose machine learning framework for predicting properties of inorganic materials. npj Comput. Mater. 2016, 2, 16028. [Google Scholar] [CrossRef] [Scilit]
  33. Ward, L.; Dunn, A.; Faghaninia, A.; Zimmermann, N.E.; Bajaj, S.; Wang, Q.; Montoya, J.; Chen, J.; Bystrom, K.; Dylla, M.; et al. Matminer: An open source toolkit for materials data mining. Comput. Mater. Sci. 2018, 152, 60–69. [Google Scholar] [CrossRef] [Scilit]
  34. Xu, P.; Ji, X.; Li, M.; Lu, W. Small data machine learning in materials science. npj Comput. Mater. 2023, 9, 42. [Google Scholar] [CrossRef] [Scilit]
  35. Dunn, A.; Wang, Q.; Ganose, A.; Dopp, D.; Jain, A. Benchmarking materials property prediction methods: The Matbench test set and Automatminer reference algorithm. npj Comput. Mater. 2020, 6, 138. [Google Scholar] [CrossRef] [Scilit]
  36. Sutton, C.; Boley, M.; Ghiringhelli, L.M.; Rupp, M.; Vreeken, J.; Scheffler, M. Identifying domains of applicability of machine learning models for materials science. Nat. Commun. 2020, 11, 4428. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Li, K.; DeCost, B.; Choudhary, K.; Greenwood, M.; Hattrick-Simpers, J. A critical examination of robustness and generalizability of machine learning prediction of materials properties. npj Comput. Mater. 2023, 9, 55. [Google Scholar] [CrossRef] [Scilit]
  38. Tavazza, F.; DeCost, B.; Choudhary, K. Uncertainty prediction for machine learning models of material properties. ACS Omega 2021, 6, 32431–32440. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Zhong, X.; Gallagher, B.; Liu, S.; Kailkhura, B.; Hiszpanski, A.; Han, T.Y.J. Explainable machine learning in materials science. npj Comput. Mater. 2022, 8, 204. [Google Scholar] [CrossRef] [Scilit]
  40. Shen, C.; Wang, C.; Wei, X.; Li, Y.; van der Zwaag, S.; Xu, W. Physical metallurgy-guided machine learning and artificial intelligent design of ultrahigh-strength stainless steel. Acta Mater. 2019, 179, 201–214. [Google Scholar] [CrossRef] [Scilit]
  41. Huang, X.; Wang, H.; Xue, W.; Ullah, A.; Xiang, S.; Huang, H.; Meng, L.; Ma, G.; Zhang, G. A combined machine learning model for the prediction of time-temperature-transformation diagrams of high-alloy steels. J. Alloys Compd. 2020, 823, 153694. [Google Scholar] [CrossRef] [Scilit]
  42. Geng, X.; Mao, X.; Wu, H.H.; Wang, S.; Xue, W.; Zhang, G.; Ullah, A.; Wang, H. A hybrid machine learning model for predicting continuous cooling transformation diagrams in welding heat-affected zone of low alloy steels. J. Mater. Sci. Technol. 2022, 107, 207–215. [Google Scholar] [CrossRef] [Scilit]
  43. Geng, X.; Wang, S.; Ullah, A.; Wu, G.; Wang, H. Prediction of hardenability curves for non-boron steels via a combined machine learning model. Materials 2022, 15, 3127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Geng, X.; Cheng, Z.; Wang, S.; Peng, C.; Ullah, A.; Wang, H.; Wu, G. A data-driven machine learning approach to predict the hardenability curve of boron steels and assist alloy design. J. Mater. Sci. 2022, 57, 10755–10768. [Google Scholar] [CrossRef] [Scilit]
  45. Cao, Y.; Zhang, C.; Tang, S.; Wu, S.; Zhou, X.; Cao, G.; Luo, D.; Wang, H.; Hedström, P.; Liu, Z. Machine learning to predict phase transformation products and their morphologies–application in design of lean high strength steel. Mater. Des. 2025, 258, 114642. [Google Scholar] [CrossRef] [Scilit]
  46. Azimi, S.M.; Britz, D.; Engstler, M.; Fritz, M.; Mücklich, F. Advanced steel microstructural classification by deep learning methods. Sci. Rep. 2018, 8, 2128. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Shen, C.; Wang, C.; Huang, M.; Xu, N.; van der Zwaag, S.; Xu, W. A generic high-throughput microstructure classification and quantification method for regular SEM images of complex steel microstructures combining EBSD labeling and deep learning. J. Mater. Sci. Technol. 2021, 93, 191–204. [Google Scholar] [CrossRef] [Scilit]
  48. Stuckner, J.; Harder, B.; Smith, T.M. Microstructure segmentation with deep learning encoders pre-trained on a large microscopy dataset. npj Comput. Mater. 2022, 8, 200. [Google Scholar] [CrossRef] [Scilit]
  49. Liu, J.; Cao, G.; Wang, H.; Cui, C.; Liu, Z. Development of intelligent methodologies perceiving microstructure and mechanical properties of hot rolled steels. Measurement 2023, 221, 113526. [Google Scholar] [CrossRef] [Scilit]
  50. Guo, S.; Yu, J.; Liu, X.; Wang, C.; Jiang, Q. A predicting model for properties of steel using the industrial big data based on machine learning. Comput. Mater. Sci. 2019, 160, 95–104. [Google Scholar] [CrossRef] [Scilit]
  51. Jiang, X.; Jia, B.; Zhang, G.; Zhang, C.; Wang, X.; Zhang, R.; Yin, H.; Qu, X.; Song, Y.; Su, L.; et al. A strategy combining machine learning and multiscale calculation to predict tensile strength for pearlitic steel wires with industrial data. Scr. Mater. 2020, 186, 272–277. [Google Scholar] [CrossRef] [Scilit]
  52. Cui, C.; Cao, G.; Cao, Y.; Liu, J.; Dong, Z.; Wu, S.; Liu, Z. Physical metallurgy guided deep learning for yield strength of hot-rolled steel based on the small labeled dataset. Mater. Des. 2022, 223, 111269. [Google Scholar] [CrossRef] [Scilit]
  53. Wu, S.W.; Yang, J.; Cao, G.M. Prediction of the Charpy V-notch impact energy of low carbon steel using a shallow neural network and deep learning. Int. J. Miner. Metall. Mater. 2021, 28, 1309–1320. [Google Scholar] [CrossRef] [Scilit]
  54. Shang, C.; Wang, C.; Wu, H.; Liu, W.; Chen, Y.; Pan, G.; Wang, S.; Wu, G.; Gao, J.; Zhao, H.; et al. Improved data-driven performance of Charpy impact toughness via literature-assisted production data in pipeline steel. Sci. China Technol. Sci. 2023, 66, 2069–2079. [Google Scholar] [CrossRef] [Scilit]
  55. Shaheen, M.A.; Presswood, R.; Afshan, S. Application of Machine Learning to predict the mechanical properties of high strength steel at elevated temperatures based on the chemical composition. In Structures; Elsevier: Amsterdam, The Netherlands, 2023; Volume 52, pp. 17–29. [Google Scholar]
  56. Peng, J.; Yamamoto, Y.; Hawk, J.A.; Lara-Curzio, E.; Shin, D. Coupling physics in machine learning to predict properties of high-temperatures alloys. npj Comput. Mater. 2020, 6, 141. [Google Scholar] [CrossRef] [Scilit]
  57. Diao, Y.; Yan, L.; Gao, K. A strategy assisted machine learning to process multi-objective optimization for improving mechanical properties of carbon steels. J. Mater. Sci. Technol. 2022, 109, 86–93. [Google Scholar] [CrossRef] [Scilit]
  58. Lee, J.Y.; Kim, M.; Lee, Y.K. Design of high strength medium-Mn steel using machine learning. Mater. Sci. Eng. A 2022, 843, 143148. [Google Scholar] [CrossRef] [Scilit]
  59. Lee, J.Y.; Kim, S.H.; Jeong, H.B.; Lee, K.; Cho, K.; Lee, Y.K. Inverse design of high-strength medium-Mn steel using a machine learning--aided genetic algorithm approach. J. Mater. Res. Technol. 2024, 33, 2672–2682. [Google Scholar] [CrossRef] [Scilit]
  60. Wang, Z.L.; Adachi, Y. Property prediction and properties-to-microstructure inverse analysis of steels by a machine-learning approach. Mater. Sci. Eng. A 2019, 744, 661–670. [Google Scholar] [CrossRef] [Scilit]
  61. Pei, Z.; Rozman, K.A.; Doğan, Ö.N.; Wen, Y.; Gao, N.; Holm, E.A.; Hawk, J.A.; Alman, D.E.; Gao, M.C. Machine-learning microstructure for inverse material design. Adv. Sci. 2021, 8, 2101207. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Kusampudi, N.; Diehl, M. Inverse design of dual-phase steel microstructures using generative machine learning model and Bayesian optimization. Int. J. Plast. 2023, 171, 103776. [Google Scholar] [CrossRef] [Scilit]
  63. Lertkiatpeeti, K.; Janya-Anurak, C.; Uthaisangsuk, V. Effects of spatial microstructure characteristics on mechanical properties of dual phase steel by inverse analysis and machine learning approach. Comput. Mater. Sci. 2024, 245, 113311. [Google Scholar] [CrossRef] [Scilit]
  64. Allen, L.; Gill, A.; Smith, A.; Hill, D.; Moghadam, P.Z.; Cordiner, J. Development of a machine learning framework to determine optimal alloy composition based on steel hardenability prediction. Digit. Chem. Eng. 2023, 9, 100118. [Google Scholar] [CrossRef] [Scilit]
  65. Li, X.; Zheng, M.; Yang, X.; Chen, P.; Ding, W. A property-oriented design strategy of high-strength ductile RAFM steels based on machine learning. Mater. Sci. Eng. A 2022, 840, 142891. [Google Scholar] [CrossRef] [Scilit]
  66. Liu, C.; Wang, X.; Cai, W.; Yang, J.; Su, H. Optimal design of the austenitic stainless-steel composition based on machine learning and genetic algorithm. Materials 2023, 16, 5633. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. Wang, J.; Lu, Y.; Wang, X.; Liang, S.; Xiong, J.; Zhen, L.; Liu, L. Target-driven design of high strength yet corrosion resistant medium Mn steel via interpretable machine-learning. Mater. Des. 2025, 260, 115217. [Google Scholar] [CrossRef] [Scilit]
  68. Zhou, X.; Xu, J.; Meng, L.; Wang, W.; Zhang, N.; Jiang, L. Machine-Learning-Assisted Composition Design for High-Yield-Strength TWIP Steel. Metals 2024, 14, 952. [Google Scholar] [CrossRef] [Scilit]
  69. Pei, Z.; Zhang, D.; Zhi, Y.; Yang, T.; Jin, L.; Fu, D.; Cheng, X.; Terryn, H.A.; Mol, J.M.; Li, X. Towards understanding and prediction of atmospheric corrosion of an Fe/Cu corrosion sensor via machine learning. Corros. Sci. 2020, 170, 108697. [Google Scholar] [CrossRef] [Scilit]
  70. Yan, L.; Diao, Y.; Lang, Z.; Gao, K. Corrosion rate prediction and influencing factors evaluation of low-alloy steels in marine atmosphere using machine learning approach. Sci. Technol. Adv. Mater. 2020, 21, 359–370. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Song, X.; Wang, K.; Zhou, L.; Chen, Y.; Ren, K.; Wang, J.; Zhang, C. Multi-factor mining and corrosion rate prediction model construction of carbon steel under dynamic atmospheric corrosion environment. Eng. Fail. Anal. 2022, 134, 105987. [Google Scholar] [CrossRef] [Scilit]
  72. Liu, M.; Li, W. Prediction and analysis of corrosion rate of 3C steel using interpretable machine learning methods. Mater. Today Commun. 2023, 35, 106408. [Google Scholar] [CrossRef] [Scilit]
  73. Dong, Z.; Ding, L.; Meng, Z.; Xu, K.; Mao, Y.; Chen, X.; Ye, H.; Poursaee, A. Machine learning-based corrosion rate prediction of steel embedded in soil. Sci. Rep. 2024, 14, 18194. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Ji, H.; Ye, H. Machine learning prediction of corrosion rate of steel in carbonated cementitious mortars. Cem. Concr. Compos. 2023, 143, 105256. [Google Scholar] [CrossRef] [Scilit]
  75. Muthukrishnan, R.; Balogun, Y.; Rajendran, V.; Prathuru, A.; Hossain, M.; Faisal, N.H. Machine learning approach to investigate high temperature corrosion of critical infrastructure materials. High Temp. Corros. Mater. 2024, 101, 309–331. [Google Scholar] [CrossRef] [Scilit]
  76. Tawk, M.; Mustapha, S.; Trad, F.; Saad, G. Service Life Prediction in Pipelines Using Machine Learning Techniques. Int. J. Comput. Intell. Syst. 2025, 18, 319. [Google Scholar] [CrossRef] [Scilit]
  77. Feng, Z.; Jiang, Z.; Yan, C. Prediction of corrosion lifetime of marine environment steel pipe piles based on physics-informed neural network. Discov. Appl. Sci. 2026, 8, 136. [Google Scholar] [CrossRef] [Scilit]
  78. Naser, M.Z. Deriving temperature-dependent material models for structural steel through artificial intelligence. Constr. Build. Mater. 2018, 191, 56–68. [Google Scholar] [CrossRef] [Scilit]
  79. Yazici, C.; Domínguez-Gutiérrez, F.J. Machine learning techniques for estimating high–temperature mechanical behavior of high strength steels. Results Eng. 2025, 25, 104242. [Google Scholar] [CrossRef] [Scilit]
  80. Zhao, Z. Steel columns under fire—A neural network based strength model. Adv. Eng. Softw. 2006, 37, 97–105. [Google Scholar] [CrossRef] [Scilit]
  81. Possidente, L.; Couto, C. Explained fire resistance machine learning models for compressed steel members of trusses and bracing systems. Eng. Appl. Artif. Intell. 2025, 139, 109571. [Google Scholar] [CrossRef] [Scilit]
  82. Fu, F. Fire induced progressive collapse potential assessment of steel framed buildings using machine learning. J. Constr. Steel Res. 2020, 166, 105918. [Google Scholar] [CrossRef] [Scilit]
  83. Kodur, V.K.; Naser, M.Z. Classifying bridges for the risk of fire hazard via competitive machine learning. Adv. Bridge Eng. 2021, 2, 2. [Google Scholar] [CrossRef] [Scilit]
  84. Naser, M.Z.; Kodur, V.; Thai, H.T.; Hawileh, R.; Abdalla, J.; Degtyarev, V.V. StructuresNet and FireNet: Benchmarking databases and machine learning algorithms in structural and fire engineering domains. J. Build. Eng. 2021, 44, 102977. [Google Scholar] [CrossRef] [Scilit]
  85. Çiftçioğlu, A.Ö.; Naser, M.Z. Fire resistance evaluation through synthetic fire tests and generative adversarial networks. Front. Struct. Civ. Eng. 2024, 18, 587–614. [Google Scholar] [CrossRef] [Scilit]
  86. Agrawal, A.; Deshpande, P.D.; Cecen, A.; Basavarsu, G.P.; Choudhary, A.N.; Kalidindi, S.R. Exploration of data science techniques to predict fatigue strength of steel from composition and processing parameters. Integr. Mater. Manuf. Innov. 2014, 3, 90–108. [Google Scholar] [CrossRef] [Scilit]
  87. Agrawal, A.; Choudhary, A. An online tool for predicting fatigue strength of steel alloys based on ensemble data mining. Int. J. Fatigue 2018, 113, 389–400. [Google Scholar] [CrossRef] [Scilit]
  88. Arvanitis, K.; Nikolakopoulos, P.; Pavlou, D.; Farmanbar, M. Machine learning-based fatigue lifetime prediction of structural steels. Alex. Eng. J. 2025, 125, 55–66. [Google Scholar] [CrossRef] [Scilit]
  89. Feng, C.; Su, M.; Xu, L.; Zhao, L.; Han, Y. Estimation of fatigue life of welded structures incorporating importance analysis of influence factors: A data-driven approach. Eng. Fract. Mech. 2023, 281, 109103. [Google Scholar] [CrossRef] [Scilit]
  90. Schubnell, J.; Fliegener, S.; Rosenberger, J.; Feth, S.; Braun, M.; Beiler, M.; Baumgartner, J. Data-driven fatigue assessment of welded steel joints based on transfer learning. Weld. World 2025, 69, 2223–2238. [Google Scholar] [CrossRef] [Scilit]
  91. Feng, C.; Su, M.; Xu, L.; Zhao, L.; Han, Y.; Peng, C. A novel generalization ability-enhanced approach for corrosion fatigue life prediction of marine welded structures. Int. J. Fatigue 2023, 166, 107222. [Google Scholar] [CrossRef] [Scilit]
  92. Jiang, L.; Hu, Y.; Liu, Y.; Zhang, X.; Kang, G.; Kan, Q. Physics-informed machine learning for low-cycle fatigue life prediction of 316 stainless steels. Int. J. Fatigue 2024, 182, 108187. [Google Scholar] [CrossRef] [Scilit]
  93. Kamble, R.G.; Raykar, N.R.; Jadhav, D.N. Machine learning approach to predict fatigue crack growth. Mater. Today Proc. 2021, 38, 2506–2511. [Google Scholar] [CrossRef] [Scilit]
  94. Long, X.; Yu, M.; Liao, W.; Jiang, C. A deep learning-based fatigue crack growth rate measurement method using mobile phones. Int. J. Fatigue 2023, 167, 107327. [Google Scholar] [CrossRef] [Scilit]
  95. Zhang, L.; Wang, Z.; Wang, L.; Zhang, Z.; Chen, X.; Meng, L. Machine learning-based real-time visible fatigue crack growth detection. Digit. Commun. Netw. 2021, 7, 551–558. [Google Scholar] [CrossRef] [Scilit]
  96. Hu, J.; Ma, K.; Zhang, Z.; Zhang, R.; Zheng, J. Machine learning-based prediction of hydrogen-assisted fatigue crack growth rate in Cr–Mo steel. Int. J. Hydrogen Energy 2025, 122, 1–11. [Google Scholar] [CrossRef] [Scilit]
  97. Himanen, L.; Geurts, A.; Foster, A.S.; Rinke, P. Data-driven materials science: Status, challenges, and perspectives. Adv. Sci. 2019, 6, 1900808. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  98. Raccuglia, P.; Elbert, K.C.; Adler, P.D.; Falk, C.; Wenny, M.B.; Mollo, A.; Zeller, M.; Friedler, S.A.; Schrier, J.; Norquist, A.J. Machine-learning-assisted materials discovery using failed experiments. Nature 2016, 533, 73–76. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Butler, K.T.; Choudhary, K.; Csanyi, G.; Ganose, A.M.; Kalinin, S.V.; Morgan, D. Setting standards for data driven materials science. npj Comput. Mater. 2024, 10, 231. [Google Scholar] [CrossRef] [Scilit]
  100. Scheffler, M.; Aeschlimann, M.; Albrecht, M.; Bereau, T.; Bungartz, H.J.; Felser, C.; Greiner, M.; Groß, A.; Koch, C.T.; Kremer, K.; et al. FAIR data enabling new horizons for materials research. Nature 2022, 604, 635–642. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  101. Ward, L.; Aykol, M.; Blaiszik, B.; Foster, I.; Meredig, B.; Saal, J.; Suram, S. Strategies for accelerating the adoption of materials informatics. MRS Bull. 2018, 43, 683–689. [Google Scholar] [CrossRef] [Scilit]
  102. Vasudevan, R.K.; Choudhary, K.; Mehta, A.; Smith, R.; Kusne, G.; Tavazza, F.; Vlcek, L.; Ziatdinov, M.; Kalinin, S.V.; Hattrick-Simpers, J. Materials science in the artificial intelligence age: High-throughput library generation, machine learning, and a pathway from correlations to the underpinning physics. MRS Commun. 2019, 9, 821–838. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  103. Tran, K.; Neiswanger, W.; Yoon, J.; Zhang, Q.; Xing, E.; Ulissi, Z.W. Methods for comparing uncertainty quantifications for material property predictions. Mach. Learn. Sci. Technol. 2020, 1, 025006. [Google Scholar] [CrossRef] [Scilit]
  104. Karniadakis, G.E.; Kevrekidis, I.G.; Lu, L.; Perdikaris, P.; Wang, S.; Yang, L. Physics-informed machine learning. Nat. Rev. Phys. 2021, 3, 422–440. [Google Scholar] [CrossRef] [Scilit]
  105. Willard, J.; Jia, X.; Xu, S.; Steinbach, M.; Kumar, V. Integrating scientific knowledge with machine learning for engineering and environmental systems. ACM Comput. Surv. 2022, 55, 66. [Google Scholar] [CrossRef] [Scilit]
  106. Gubernatis, J.E.; Lookman, T.J.P.R.M. Machine learning in materials design and discovery: Examples from the present and suggestions for the future. Phys. Rev. Mater. 2018, 2, 120301. [Google Scholar] [CrossRef] [Scilit]
  107. Lookman, T.; Balachandran, P.V.; Xue, D.; Yuan, R. Active learning in materials science with emphasis on adaptive sampling using uncertainties for targeted design. npj Comput. Mater. 2019, 5, 21. [Google Scholar] [CrossRef] [Scilit]
  108. Xue, D.; Balachandran, P.V.; Hogden, J.; Theiler, J.; Xue, D.; Lookman, T. Accelerated search for materials with targeted properties by adaptive design. Nat. Commun. 2016, 7, 11241. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  109. Kusne, A.G.; Yu, H.; Wu, C.; Zhang, H.; Hattrick-Simpers, J.; DeCost, B.; Sarker, S.; Oses, C.; Toher, C.; Curtarolo, S.; et al. On-the-fly closed-loop materials discovery via Bayesian active learning. Nat. Commun. 2020, 11, 5966. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  110. Kalidindi, S.R.; Buzzy, M.; Boyce, B.L.; Dingreville, R. Digital twins for materials. Front. Mater. 2022, 9, 818535. [Google Scholar] [CrossRef] [Scilit]
  111. Artificial Intelligence Risk Management Framework (AI RMF 1.0). 2023. Available online: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (accessed on 19 August 2026).
Figure 1. Conceptual framework of multisource data integration and machine-learning applications for structural steels.
Figure 1. Conceptual framework of multisource data integration and machine-learning applications for structural steels.
Materials 19 03612 g001
Figure 2. Machine-learning workflow and trustworthy evaluation framework for structural steels [31]. The schematic elements (e.g., feature variables, sample matrix, and candidate points) are used to illustrate general machine-learning procedures rather than specific datasets. Ellipses indicate multiple unshown variables or samples, and different colors are used only for visual distinction.
Figure 2. Machine-learning workflow and trustworthy evaluation framework for structural steels [31]. The schematic elements (e.g., feature variables, sample matrix, and candidate points) are used to illustrate general machine-learning procedures rather than specific datasets. Ellipses indicate multiple unshown variables or samples, and different colors are used only for visual distinction.
Materials 19 03612 g002
Figure 3. Machine-learning-assisted framework for phase-transformation prediction, alloy-lean design, and process optimization.
Figure 3. Machine-learning-assisted framework for phase-transformation prediction, alloy-lean design, and process optimization.
Materials 19 03612 g003
Figure 4. Transfer-learning-assisted deep-learning segmentation and quantitative analysis of complex microstructures [48].
Figure 4. Transfer-learning-assisted deep-learning segmentation and quantitative analysis of complex microstructures [48].
Materials 19 03612 g004
Figure 5. Experimental and synthetic physics-based features used to predict the yield strength of 9–12Cr steels [56].
Figure 5. Experimental and synthetic physics-based features used to predict the yield strength of 9–12Cr steels [56].
Materials 19 03612 g005
Figure 6. Materials-design approach based on microstructural images and a variational autoencoder [61]. (a) General procedure of material design with and without artificial intelligence (AI). (b) 32 sub-images of 128 × 128 are randomly chopped from one representative scanning electron microscopy (SEM) image that has a martensitic microstructure. The bottom black bar is excluded from image chopping. (The black squares do not represent the real sizes of sub-images). (c) The schematic for the machine-learning model (i.e., VAE) used in this study. The model consists of three sub-models, i.e., the encoder model, the decoder model, and a regression model.
Figure 6. Materials-design approach based on microstructural images and a variational autoencoder [61]. (a) General procedure of material design with and without artificial intelligence (AI). (b) 32 sub-images of 128 × 128 are randomly chopped from one representative scanning electron microscopy (SEM) image that has a martensitic microstructure. The bottom black bar is excluded from image chopping. (The black squares do not represent the real sizes of sub-images). (c) The schematic for the machine-learning model (i.e., VAE) used in this study. The model consists of three sub-models, i.e., the encoder model, the decoder model, and a regression model.
Materials 19 03612 g006
Figure 7. Machine-learning-assisted composition-design workflow for high-yield-strength TWIP steels [68].
Figure 7. Machine-learning-assisted composition-design workflow for high-yield-strength TWIP steels [68].
Materials 19 03612 g007
Figure 8. Machine-learning framework for evaluating corrosion degradation and service life of structural steels.
Figure 8. Machine-learning framework for evaluating corrosion degradation and service life of structural steels.
Materials 19 03612 g008
Figure 9. Measured and machine-learning-predicted long-term corrosion rates of low-alloy steels in different marine atmospheric environments [70]. (a) illustration of three different atmospheric exposure sites; the measured corrosion rate (star icons, three parallel samples) vs. predicted corrosion rate (bar plot) of typical low-alloy steel SPA-H, SMA490 and SM490A after (b) one and (c) ten years of exposure.
Figure 9. Measured and machine-learning-predicted long-term corrosion rates of low-alloy steels in different marine atmospheric environments [70]. (a) illustration of three different atmospheric exposure sites; the measured corrosion rate (star icons, three parallel samples) vs. predicted corrosion rate (bar plot) of typical low-alloy steel SPA-H, SMA490 and SM490A after (b) one and (c) ten years of exposure.
Materials 19 03612 g009
Figure 10. Multiscale machine-learning framework for evaluating the fire and extreme-temperature performance of structural steels.
Figure 10. Multiscale machine-learning framework for evaluating the fire and extreme-temperature performance of structural steels.
Materials 19 03612 g010
Figure 11. Bridge fire-risk assessment workflow based on competing machine-learning algorithms [83].
Figure 11. Bridge fire-risk assessment workflow based on competing machine-learning algorithms [83].
Materials 19 03612 g011
Figure 12. General workflow for machine-learning-based fatigue-life prediction of structural steels [88].
Figure 12. General workflow for machine-learning-based fatigue-life prediction of structural steels [88].
Materials 19 03612 g012
Figure 13. Integrated machine-learning framework for fatigue-life prediction, crack-growth monitoring, and remaining-life assessment of structural steels.
Figure 13. Integrated machine-learning framework for fatigue-life prediction, crack-growth monitoring, and remaining-life assessment of structural steels.
Materials 19 03612 g013
Figure 14. Roadmap for trustworthy machine learning in structural-steel design, manufacturing, and in-service management.
Figure 14. Roadmap for trustworthy machine learning in structural-steel design, manufacturing, and in-service management.
Materials 19 03612 g014
Table 1. Comparison of the present review with representative previous reviews on machine learning in materials and structural engineering.
Table 1. Comparison of the present review with representative previous reviews on machine learning in materials and structural engineering.
Review CategoryMain Material or Structural ScopeMajor Data ModalitiesMaterials Design and Property PredictionService Degradation and Life AssessmentTrustworthiness and Engineering Validation
Alloy and metallic-material ML reviews [4,5,6]Broad alloy and metallic systemsComposition, computation, tabular dataExtensiveLimitedPartially addressed
Microstructure and inclusion reviews [2,13]Steels and metallic materialsMicroscopy and image dataSpecializedLimitedLimited
Constitutive and fatigue reviews [14,15,16]Metallic materials and steelsMechanical and fatigue dataProperty-specificPrimarily fatigue-relatedPartially addressed
Building-material ML reviews [17]Concrete and functional construction materialsMixedBroadPartialLimited
Steel-structure AI reviews [18]Members and structural systemsStructural, numerical, and monitoring dataLimited at the material scaleStructural-performance orientedPartially addressed
Steel manufacturing and performance reviews [19,20,21,22]Steel production and materialsIndustrial and process dataExtensiveLimitedPartially addressed
Present reviewCivil-infrastructure structural steels, with explicitly classified adjacent and transfer evidenceTabular, imaging, three-dimensional, simulation, sensor, and service dataIntegrated from composition and process design to microstructure and property predictionCorrosion, fire, fatigue, fracture, and remaining lifeData independence, physical consistency, interpretability, uncertainty, applicability domain, and external validation
Table 2. Major data types used in machine learning for structural steels.
Table 2. Major data types used in machine learning for structural steels.
Data TypeRepresentative Data Scale and Independent UnitEssential MetadataAccessibilityRecommended Partitioning UnitPrincipal LimitationRef.
Composition and processing dataLarge numbers of records may originate from substantially fewer independent heats or production campaignsSteel grade, heat number, product form, thickness, complete processing route, equipment stateMainly proprietary; some literature and public datasetsHeat, steel grade, production campaign, or chronological periodStrong collinearity, process drift, and incomplete processing histories[19,20,21,22]
Mechanical and durability dataUsually limited by the number of independently tested specimens, conditions, or exposure seriesSpecimen geometry, sampling orientation, test standard, temperature, strain/loading rate, environment, run-out/censoring statusLiterature, laboratory databases, selected public repositoriesSpecimen, heat, laboratory, or exposure conditionHigh testing cost, censoring, and scarcity of failure or boundary-condition data[19,20]
Microstructural imagesImage or patch counts may substantially exceed the number of original specimens; e.g., 1705 SEM images and 8909 annotated regions in Aachen–HeerlenSpecimen ID, imaging modality, magnification, pixel size, preparation, field of view, labeling protocolSeveral open repositories availableOriginal specimen or independently acquired field of viewLabel uncertainty, imaging-scale dependence, and patch-level leakage[2,23,24]
Defect and three-dimensional dataMultiple defects may originate from a small number of independently tested volumes or specimensSpecimen ID, voxel size, reconstruction settings, detection threshold, loading history, defect definitionLimited but increasing open availabilityOriginal specimen or scanned volumeHigh acquisition cost and sensitivity to reconstruction and thresholding[13,25]
Industrial time-series and sensor dataVery large numbers of time points may originate from relatively few production campaignsSensor type, sampling frequency, calibration state, equipment condition, heat/campaign ID, maintenance eventsPredominantly proprietaryProduction campaign, heat, or chronological blockAutocorrelation, sensor drift, operating-condition change, and leakage[20,21,22]
Numerical simulation dataPotentially large parameter-space coverage; independent unit is a simulated physical state rather than an experimental specimenModel fidelity, governing equations, material parameters, boundary conditions, mesh/resolution, calibration and uncertaintyOften reproducible when models and inputs are availableParameter-space region or physical scenarioModel-form discrepancy, parameter uncertainty, and simulation-to-experiment bias[3,20,28]
Table 3. Trustworthy evaluation framework for machine-learning models of structural steels.
Table 3. Trustworthy evaluation framework for machine-learning models of structural steels.
Evaluation DimensionCore QuestionRecommended Evaluation ApproachRef.
Predictive accuracyCan the model accurately describe known samples?R2, MAE, RMSE, F1 score, AUC, and region-specific errors[19,35]
Physical consistencyAre the predicted trends consistent with metallurgical and mechanical principles?Verification against boundary conditions, variable trends, and physics-based models[3,31,39]
GeneralizationCan the model be transferred to new heats, steel grades, and service conditions?Leave-one-heat-out, leave-one-grade-out, cross-manufacturer, and external validation[36,37]
UncertaintyIs an individual prediction sufficiently reliable?Prediction intervals, coverage probability, and calibration error[38]
InterpretabilityDoes the model identify physically reasonable controlling factors?SHAP analysis, sensitivity analysis, and partial-dependence plots[39]
ReproducibilityCan the data, model, and results be independently reproduced?Public availability of data, code, metadata, and partitioning protocols[29,30,35]
Table 4. Representative applications of machine learning to composition, processing, and microstructure design of structural steels.
Table 4. Representative applications of machine learning to composition, processing, and microstructure design of structural steels.
Material or Application ScopeData and VariablesModels and ValidationKey Findings, Validation Evidence, and Applicability LimitationsRef.
Ultrahigh-strength stainless steel as a methodological transfer caseSmall literature-derived and experimental dataset; composition, aging conditions, and precipitation-related physical variables used to predict hardness and feasibilityPhysics-guided SVM, classifier, and GA; cross-validation and independent alloy fabricationPhysics-based features excluded infeasible candidates; the material is not a conventional civil-engineering structural steel[40]
Phase transformations in high- and low-alloy steelsLiterature and experimental TTT and CCT data; composition and cooling conditions used to predict transformation products, temperatures, and hardnessEnsemble regression, multilayer perceptron (MLP), RF, and symbolic regression; test-set evaluation and comparison with conventional modelsAccelerated prediction of transformation diagrams; limited consistency across databases[41,42]
Hardenability of non-boron and boron steelsJominy-test data; composition and distance from the quenched end used to predict hardness profiles and optimize compositionRF, k-nearest neighbors (kNN), and combined models; cross-validation and validation of candidate steelsSuitable for curve-valued outputs; performance depends on the coverage of compositionally similar steels[43,44]
Complex steel microstructuresSEM and EBSD images with limited numbers of original specimens; image-based phase classification, segmentation, and quantificationFully convolutional neural network (FCNN), U-Net, and other deep segmentation models; specimen-level partitioning and cross-grade testingEBSD-derived labels improved physical fidelity; performance remained sensitive to specimen preparation and magnification[46,47,48,49]
Table 5. Representative machine-learning studies for mechanical-property prediction of structural steels.
Table 5. Representative machine-learning studies for mechanical-property prediction of structural steels.
PropertyData BasisML Approach and ValidationMain LimitationRef.
Tensile propertiesIndustrial production dataTree-based/ensemble models; group-level validationIndependent heats should be clarified[51,52]
Tensile properties63,137 recordsRF/regression; internal validationRecords may not represent independent heats[50]
Charpy impact energyLiterature and industrial dataRegression/ensemble; test-set validationTemperature and specimen conditions vary[53,54]
Elevated-temperature propertiesHigh-temperature experimentsRegression/ensemble; experimental validationHeating and loading histories vary[55,56]
Table 6. Principal machine-learning tasks in structural-steel durability and service-life assessment.
Table 6. Principal machine-learning tasks in structural-steel durability and service-life assessment.
TaskMain TargetTypical Scale
Corrosion-rate predictionCorrosion rate/currentMaterial
Localized-damage predictionPit depth/morphologyLocal surface
Elevated-temperature degradationRetained propertiesMaterial
Fire-resistance predictionMember resistanceComponent
Fatigue-strength predictionFatigue strengthMaterial/joint
Total-life predictionFatigue lifeMaterial/component
Crack-growth monitoringCrack length/growth rateLocal/component
Remaining-life estimationResidual service lifeComponent/structure
Table 7. Representative machine-learning studies for corrosion and environmental degradation of structural steels.
Table 7. Representative machine-learning studies for corrosion and environmental degradation of structural steels.
Exposure ClassStudyTarget (Unit)Exposure/TimeModel and ValidationLocalized DamageRef.
AtmosphericCarbon steel sensorInstantaneous galvanic current (μA)34 dRF; time-series testingNo[69]
Marine atmosphereLow-alloy steelsCorrosion rate (μm·a−1)1–10 yrRF; train/test splitNo[70]
Dynamic atmosphereCarbon steelCorrosion rate (μm·a−1)Dynamic vehicle exposureGA-SVR; hold-out testingNo[71]
Seawater3C steelCorrosion current density (μA·cm−2)Multiple seawater conditionsANN; unseen test dataNo[71]
Soil/undergroundSteelCorrosion current density (A·m−2)0–415 dRF; random 80/20 splitNo[73]
Steel–concreteSteel in mortarCorrosion current density (μA·cm−2)Laboratory exposureSVR; cross-validationNo[74]
High-temperature corrosionStructural/coating materialsCorrosion rate (study-specific units)Heterogeneous literature dataRF; internal validationNot resolved[75]
Pipeline/service lifePipeline systemsTime to failure (source-defined time unit)Historical failure recordsExtra Trees; CV + test setNot resolved[76]
Marine pipe pilesSteel pipe pilesCorrosion depth (mm); service life (yr)Accelerated test + long-term extrapolationPINN; experimental validationNo; uniform corrosion assumption[77]
Note: “Localized damage” indicates whether the machine-learning target explicitly represents spatially concentrated corrosion such as maximum pit depth or pit morphology. “No” indicates that the model primarily predicts an average or electrochemical corrosion quantity. “Not resolved” indicates that localized damage was not explicitly represented in the reviewed model.
Table 8. Representative applications of machine learning to elevated-temperature and fire-performance assessment of structural steels.
Table 8. Representative applications of machine learning to elevated-temperature and fire-performance assessment of structural steels.
Scale or Application ScopeData and VariablesModels and ValidationKey Findings and LimitationsRef.
MaterialMultisource elevated-temperature material data; temperature and material parameters used to predict thermal-property and mechanical-property reductionANN combined with GA; comparison with code provisions and experimentsHeterogeneous data could be integrated, but definitional inconsistencies propagated into member-level calculations[78]
MaterialMore than 450 expanded records; temperature and plate thickness used to predict elevated-temperature stress–strain behaviorGradient boosting, RF, XGBoost, and SVR; experimental validationGradient boosting achieved an adjusted (R2 > 0.98); extrapolation beyond the calibrated temperature range requires caution[79]
MemberExperimental and analytical steel-column data; temperature, geometry, and loading variables used to predict compressive resistanceHybrid neural network; comparison with analytical equations and experimentsCapable of handling multiple interacting variables; coverage of the early dataset was limited[80]
Member21,879 GMNIA samples; cross-sectional geometry, temperature, and slenderness used to predict fire resistanceNeural network, RF, and SVM with SHAP analysis; independent test setSuitable as a rapid surrogate model; bias inherited from numerical simulation remained unresolved[81]
Structural systemMonte Carlo-generated fire scenarios; temperature and loading variables used to predict member failure and structural collapseDecision tree, KNN, and ANN; validation using numerical case studiesLinked member failure with system-level response; full-scale experimental evidence remained limited[82]
Infrastructure riskFire records for 118 bridges; structural and fire-related features used for risk classificationRF, SVM, and GAM; cross-validationClassification accuracy was approximately 70%; the dataset was heterogeneous and limited in size[83]
Benchmarking and data augmentationFireNet and synthetic test data; multiple inputs used to predict fire performanceBenchmark models, GAN, and VAE; unified benchmarking and internal validationImproved model comparability; synthetic data could not generate previously unobserved failure mechanisms[84,85]
Table 9. Representative applications of machine learning to fatigue, fracture, and remaining-life assessment of structural steels.
Table 9. Representative applications of machine learning to fatigue, fracture, and remaining-life assessment of structural steels.
Task/ScopeData CharacteristicsML ApproachValidation and Data TreatmentMain LimitationRef.
Fatigue-strength predictionNIMS steel fatigue data; composition and processing variablesRegression and ensemble modelsCross-validation; independent specimen number and run-out treatment not reportedApplicability restricted by database coverage[86,87]
Total fatigue-life predictionNIMS data and supplementary fatigue testsPolynomial regression, SVR, XGBoost, ANNTrain/test evaluation; specimen independence and censored-data treatment not fully reportedExternal validation remains limited[88]
Welded-joint S–N predictionWeld geometry, loading conditions, and physics-based descriptorsXGBoost and DCNNCross-validation; uncertainty not explicitly quantifiedTransferability across joint classes and loading spectra remains limited[89]
Welded-joint fatigue transfer22 target-domain welded-joint datasetsTransfer learningTarget-domain validation and comparison with conventional modelsPerformance depends on similarity between source and target domains[90]
Corrosion-fatigue lifeMultisource data covering stress ratio, frequency, temperature, and environmentXGBoost and attention-based DCNNCross-condition evaluation; prediction intervals not reportedValidation under realistic variable-amplitude loading remains insufficient[91]
Low-cycle fatigue, methodological transfer case316 stainless steel under different temperatures and strain ratesPhysics-informed neural networkFactor-of-two error-band evaluationPhysical constraints require reformulation for structural steels[92]
Fatigue-crack-growth predictionCarbon-steel compact-tension data; fracture-mechanics variablesRegression and KNNExperimental test-data evaluation; uncertainty not reportedExtrapolation beyond the tested crack-growth regime is uncertain[93]
Crack detection and monitoringCyclic-loading image sequences; surface crack length and growth rateFaster R-CNN and vision-based MLExperimental sequence evaluation; specimen-level independence should be preservedPrimarily limited to visible surface cracks[94,95]
Hydrogen-assisted crack growth, methodological transfer caseCr–Mo steel; hydrogen pressure, loading and material variablesGradient boosting and SHAPTrain/test evaluation within the investigated material and loading domainResults should not be generalized beyond the specified Cr–Mo steel and hydrogen conditions[96]
Note: NIMS, National Institute for Materials Science; SVR, support vector regression; ANN, artificial neural network; DCNN, deep convolutional neural network; KNN, k-nearest neighbors; R-CNN, region-based convolutional neural network; SHAP, SHapley Additive exPlanations. Where the original study did not explicitly report the number of statistically independent specimens, treatment of fatigue run-outs, or prediction uncertainty, this limitation is stated rather than inferred.
Table 10. Priority research gaps and readiness of machine learning for structural steels.
Table 10. Priority research gaps and readiness of machine learning for structural steels.
GapPriority ActionHorizonCurrent ReadinessEvidence Basis
Heterogeneous dataStandardize metadata and provenanceNear termDevelopingInconsistent datasets
Limited generalizationGrouped and external validationNear termDevelopingMainly internal validation
Uncertainty/OODCalibrated intervals and OOD detectionNear–mediumEarlyLimited uncertainty reporting
Physical consistencyPhysics-informed and multi-fidelity modelsMedium termDevelopingNarrow physical integration
Experimental costActive learning and transfer learningMedium termEarly–developingSmall high-cost datasets
Inverse designProspective candidate validationMedium termEarlyMainly limited experimental validation
Dynamic life assessmentOperational digital twinsLong termEarlyWeak cross-scale updating
Engineering deploymentGovernance and code integrationLong termEarlyLimited deployment evidence
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wei, G.; Li, M.; Cui, B.; Xiu, W.; Sarman, A.M. Machine Learning for Structural Steels: Materials Design, Property Prediction, Durability, and Future Directions. Materials 2026, 19, 3612. https://doi.org/10.3390/ma19173612

AMA Style

Wei G, Li M, Cui B, Xiu W, Sarman AM. Machine Learning for Structural Steels: Materials Design, Property Prediction, Durability, and Future Directions. Materials. 2026; 19(17):3612. https://doi.org/10.3390/ma19173612

Chicago/Turabian Style

Wei, Guomin, Minghe Li, Bo Cui, Wencui Xiu, and Asmawan Mohd Sarman. 2026. "Machine Learning for Structural Steels: Materials Design, Property Prediction, Durability, and Future Directions" Materials 19, no. 17: 3612. https://doi.org/10.3390/ma19173612

APA Style

Wei, G., Li, M., Cui, B., Xiu, W., & Sarman, A. M. (2026). Machine Learning for Structural Steels: Materials Design, Property Prediction, Durability, and Future Directions. Materials, 19(17), 3612. https://doi.org/10.3390/ma19173612

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop