Comparative Analysis of Various Supervised Machine Learning Models for the Prediction of the Outcome of the Welded Bead Bending Test
Abstract
1. Introduction
2. Materials and Methods
2.1. Welded Bead Bending Test (WBBT)
2.2. Dataset
2.3. Models
2.4. Performance Metrics
2.5. Stability as an Indicator of Generalisation in Machine Learning
2.6. Preprocessing, Computing Technology and Computational Cost
3. Results and Discussion
3.1. Classification Evaluation Metrics
3.2. Stability of the Performance
3.3. Computation Time
3.4. Evaluation
3.5. Feature Importance
4. Conclusions
- The models demonstrate disparate performances for both classes (p/n.p.) and in between training and testing datasets. In the majority of cases, the performance on the known training dataset is superior to that on the unknown testing dataset. A very precise prediction of the n.p. class, which corresponds to high Specificity, is crucial for any potential partial substitution of the WBBT.
- The most stable predictions over all three metrics are achieved by GLVQ ( up to ) and GMLVQ ( up to ). This is indicative of the excellent generalisation and robustness of these models.
- The computation time separated the models into three distinct categories: KNN and DT required approximately 0.003 s to 0.07 s for their training. BCDT, HGBC and RF required less than 1 s. GLVQ, BCRF and GMLVQ needed approximately 3 s to 15 s. The latter model exhibited a number of outliers, characterised by elevated computation times (up to 80 s), whereas the BCRF required a consistent time of 10 s to 20 s. It is imperative to acknowledge that the training of a single model is sufficient to enable prediction. This precludes the computation time from being a major performance criterion.
- BCRF and BCDT exhibited optimal model performance, demonstrating the highest single-trial Balanced Accuracy (70.3%; 70.6%, respectively). BCRF achieved superior Specificity (82.5%) with a Recall of 58%. The n.p. class is detected with a high degree of accuracy. KNN showed the lowest suitability for WBBT outcome prediction. Despite attaining the maximum Recall, this can be ascribed to overfitting, given that the Specificity was considerably diminished.
- Feature importance analysis for BCRF identified material thickness (20.17%), Z-grade (17.62%), and its standard deviation (17.46%) as the most decisive predictors. The high significance of through-thickness ductility and homogeneity suggests that these parameters are more critical for the WBBT outcome than traditionally emphasised criteria, such as CVN (4.24%).
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| BCDT | Bagging Classifier with Decision Tree Classifier as base estimator |
| BA | Balanced Accuracy |
| BCRF | Bagging Classifier with Random Forest Classifier as base estimator |
| CEWUS | Chemnitzer Werkstoff und Oberflächentechnik |
| CVN | Charpy V-Notch impact energy |
| DT | Decision Tree Classifier |
| ESF | European Social Fund |
| FN | False Negative |
| FP | False Positive |
| GLVQ | Generalized Learning Vector Quantizer |
| GMLVQ | Generalized Matrix Learning Vector Quantizer |
| HAZ | Heat-Affected Zone |
| HGBC | Histogram-based Gradient Boosting Classifier |
| KNN | k-Nearest Neighbours |
| ML | Machine Learning |
| n.p. class | Samples that did not pass the Welded Bead Bending Test |
| OES | Optical emission spectrometry |
| p class | Samples that passed the Welded Bead Bending Test |
| REC | Recall |
| RF | Random Forest Classifier |
| SEP | Stahl-Eisen-Prüfblatt |
| SPEC | Specificity |
| std. dev. | Standard deviation |
| TN | True Negative |
| TP | True Positive |
| WBBT | Welded Bead Bending Test |
| Z | Z-Grade from through-thickness tensile test |
Appendix A. Performance Metrics
Appendix A.1. 50 Trials
| Model | Training Dataset | Testing Dataset | |||||
|---|---|---|---|---|---|---|---|
| BA (%) | REC (%) | SPEC (%) | BA (%) | REC (%) | SPEC (%) | ||
| DT | Min | 56.90 | 26.97 | 17.61 | 50.41 | 23.00 | 11.11 |
| Max | 81.25 | 96.23 | 96.26 | 92.45 | 96.49 | 92.45 | |
| Mean | 66.73 | 58.69 | 74.77 | 59.58 | 58.12 | 59.58 | |
| RF | Min | 66.11 | 42.96 | 81.03 | 51.44 | 38.93 | 40.91 |
| Max | 83.32 | 72.26 | 98.27 | 66.00 | 73.31 | 82.98 | |
| Mean | 75.99 | 62.21 | 89.77 | 60.68 | 60.67 | 60.68 | |
| HGBC | Min | 82.95 | 67.86 | 90.10 | 47.18 | 63.33 | 30.00 |
| Max | 84.66 | 79.12 | 98.99 | 64.64 | 74.88 | 62.07 | |
| Mean | 83.83 | 71.21 | 96.46 | 58.23 | 69.36 | 47.11 | |
| KNN | Min | 74.31 | 90.11 | 48.92 | 47.94 | 85.95 | 2.38 |
| Max | 80.46 | 99.78 | 64.29 | 59.79 | 97.92 | 27.45 | |
| Mean | 76.95 | 96.58 | 57.32 | 53.51 | 93.44 | 13.58 | |
| BCDT | Min | 50.00 | 32.67 | 0.00 | 50.00 | 22.58 | 0.00 |
| Max | 82.76 | 100.00 | 92.05 | 70.57 | 100.00 | 91.35 | |
| Mean | 70.04 | 78.01 | 62.06 | 60.44 | 59.57 | 61.31 | |
| BCRF | Min | 70.46 | 56.80 | 80.93 | 53.86 | 53.32 | 42.42 |
| Max | 81.50 | 71.51 | 96.30 | 70.25 | 71.72 | 82.50 | |
| Mean | 75.37 | 62.60 | 88.13 | 63.46 | 60.79 | 66.13 | |
| GLVQ | Min | 50.00 | 6.04 | 0.00 | 50.00 | 6.16 | 0.00 |
| Max | 60.61 | 100.00 | 99.49 | 63.74 | 100.00 | 100.00 | |
| Mean | 55.84 | 56.04 | 55.65 | 55.81 | 56.09 | 55.53 | |
| GMLVQ | Min | 50.00 | 0.00 | 18.09 | 47.89 | 0.00 | 10.64 |
| Max | 61.33 | 92.38 | 100.00 | 65.26 | 93.82 | 100.00 | |
| Mean | 56.31 | 55.89 | 56.73 | 55.10 | 55.35 | 54.84 | |
Appendix A.2. Best Single Trials
| Model | Balanced Accuracy (%) | Recall (%) | Specificity (%) |
|---|---|---|---|
| DT | 65.57 | 63.79 | 67.35 |
| RF | 66.00 | 64.55 | 67.44 |
| HGBC | 64.64 | 67.21 | 62.07 |
| KNN | 56.78 | 89.88 | 23.68 |
| BCDT | 70.57 | 64.10 | 77.05 |
| BCRF | 70.25 | 58.00 | 82.50 |
| GLVQ | 63.74 | 49.93 | 77.55 |
| GMLVQ | 65.26 | 70.52 | 60.00 |
Appendix B. Computation Time
| Model | Min (s) | Max (s) | Mean (s) |
|---|---|---|---|
| DT | 0.003 | 0.024 | 0.008 |
| RF | 0.149 | 0.759 | 0.459 |
| HGBC | 0.030 | 0.376 | 0.153 |
| KNN | 0.005 | 0.068 | 0.021 |
| BCDT | 0.003 | 0.937 | 0.123 |
| BCRF | 5.32 | 29.74 | 14.33 |
| GLVQ | 0.126 | 4.416 | 2.875 |
| GMLVQ | 1.63 | 80.52 | 10.27 |
References
- Salvati, E.; Tognan, A.; Laurenti, L.; Pelegatti, M.; Bona, F.D. A defect-based physics-informed machine learning framework for fatigue finite life prediction in additive manufacturing. Mater. Des. 2022, 222, 111089. [Google Scholar] [CrossRef] [Scilit]
- Bruno, F.; Konstantoupoulos, G.; Rossi, E.; Fiore, G.; Charitidis, C.; Sebastiani, M.; Belforte, L.; Palumbo, M. Advanced microstructural characterization in high-strength steels via machine learning-enhanced high-speed nanoindentation and EBSD mapping. Mater. Today Commun. 2024, 39, 109192. [Google Scholar] [CrossRef] [Scilit]
- Chang, J.; Li, J.; Ye, J.; Zhang, B.; Chen, J.; Xia, Y.; Lei, J.; Carlson, T.; Loureiro, R.; Korsunsky, A.M.; et al. AI-Enabled Piezoelectric Wearable for Joint Torque Monitoring. Nano-Micro Lett. 2025, 17, 247. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Forschungsgesellschaft für Straßen- und Verkehrswesen (FGSV). Zusätzliche Technische Vertragsbedingungen und Richtlinien für Ingenieurbauten ZTV-ING; Technical report; Bundesministerium für Digitales und Verkehr (BMDV): Berlin, Germany, 2023. [Google Scholar]
- DBS 918 002-02; Technische Lieferbedingungen: Warmgewalzte Erzeugnisse aus Baustählen für den Eisenbahnbrückenbau. Technical report; Deutsche Bahn AG: Berlin, Germany, 2006.
- SEP 1390; Aufschweißbiegeversuch. Technical report; Verein Deutscher Eisenhüttenleute (VDEh): Düsseldorf, Germany, 1996.
- Stranghöner, N. Der Aufschweißbiegeversuch oder: Nichts ist beständiger als ein Provisorium. Stahlbau 2009, 78, 815–821. [Google Scholar] [CrossRef] [Scilit]
- DIN EN 1993-1-1:2025-04; Eurocode 3: Design of Steel Structures—Part 1-1: General Rules and Rules for Buildings. Technical report; Deutsches Institut für Normung (DIN): Berlin, Germany, 2025.
- Feldmann, M.; Citarelli, S.; Münstermann, S.; Könemann, M. AUBI-äquivalente Anforderungen an die Zähigkeitshochlage aus Versuchen und Schädigungssimulationen. Stahlbau 2020, 89, 1016–1026. [Google Scholar] [CrossRef] [Scilit]
- Sedlacek, G.; Höhler, S.; Dahl, W.; Kühn, B.; Langenberg, P.; Finger, M.; Floßdorf, F.J.; Schröter, F.; Hocké, A. Ersatz des Aufschweißbiegeversuchs durch äquivalente Stahlgütewahl. Stahlbau 2005, 74, 539–546. [Google Scholar] [CrossRef] [Scilit]
- DIN EN ISO 2560; Welding Consumables—Covered Electrodes for Manual Metal Arc Welding of Non-Alloy and Fine Grain Steels—Classification. Technical report; Deutsches Institut für Normung (DIN): Berlin, Germany, 2021.
- Backofen, F.; Hähnel, U.; Hahn, F.; Hockauf, M.; Kaiser, P.; Hockauf, K. Prediction of the Outcome of a Welded Bead Bending Test (WBBT) using Machine Learning (ML). In Proceedings of the 43. Vortrags- und Diskussionstagung Werkstoffprüfung 2025; Zimmermann, M., Ed.; Deutsche Gesellschaft für Materialkunde e.V. (DGM): Nordrhein-Westfalen, Germany, 2025; pp. 86–91. [Google Scholar]
- Breiman, L.; Friedman, J.H.; Olshen, R.A.; Review, C.J.S.; Gordon, A.D. Classification and Regression Trees; Wadsworth International Group: New York, NY, USA, 1984; Volume 40, p. 874. [Google Scholar]
- Quinlan, J.R. Induction of Decision Trees. Mach. Learn. 1986, 1, 81–106. [Google Scholar] [CrossRef] [Scilit]
- Cover, T.M.; Hart, P.E. Nearest Neighbor Pattern Classification. IEEE Trans. Inf. Theory 1967, 13, 21–27. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Bagging Predictors. Mach. Learn. 1996, 24, 123–140. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems 30; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 3146–3154. Available online: https://dl.acm.org/doi/10.5555/3294996.3295074 (accessed on 7 April 2026).
- Theerthagiri, P. Liver disease classification using histogram-based gradient boosting classification tree with feature selection algorithm. Biomed. Signal Process. Control 2025, 100, 107102. [Google Scholar] [CrossRef] [Scilit]
- Sato, A.; Yamada, K. Generalized Learning Vector Quantization. In Proceedings of the Advances in Neural Information Processing Systems; MIT Press: Cambridge, MA, USA, 1996; pp. 423–429. [Google Scholar]
- Schneider, P.; Schleif, F.M.; Villmann, T.; Biehl, M. Generalized Matrix Learning Vector Quantizer for the Analysis of Spectral Data. In Proceedings of the 16th European Symposium on Artificial Neural Networks (ESANN 2008), Bruges, Belgium, 23–25 April 2008; D-Side Publications: Bruges, Belgium, 2008; pp. 451–456. [Google Scholar]
- Rojek, I.; Jasiulewicz-Kaczmarek, M.; Piechowski, M.; Mikołajewski, D. The use of decision trees to identify the causes of failures in a medical enterprise—A case study. In Proceedings of the IFAC-PapersOnLine; Elsevier B.V.: Amsterdam, The Netherlands, 2024; Volume 58, pp. 133–138. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y.; Chang, Y.; Li, K. Applications of K-nearest neighbor algorithm in intelligent diagnosis of wind turbine blades damage. Renew. Energy 2023, 212, 855–864. [Google Scholar] [CrossRef] [Scilit]
- Tharwat, A.; Gaber, T.; Awad, Y.M.; Dey, N.; Hassanien, A.E. Plants identification using feature fusion technique and bagging classifier. In Proceedings of the Advances in Intelligent Systems and Computing; Springer: Berlin/Heidelberg, Germany, 2016; Volume 407, pp. 461–471. [Google Scholar] [CrossRef] [Scilit]
- Patel, R.K.; Giri, V. Feature selection and classification of mechanical fault of an induction motor using random forest classifier. Perspect. Sci. 2016, 8, 334–337. [Google Scholar] [CrossRef] [Scilit]
- Deb, C.; Nachiappan, M.R.; Elangovan, M.; Sugumaran, V. Fault Diagnosis of a Single Point Cutting Tool using Statistical Features by Random Forest Classifier. Indian J. Sci. Technol. 2016, 9, 1–8. [Google Scholar] [CrossRef] [Scilit]
- Saifudin, A.; Nabillah, U.U.; Yulianti; Desyani, T. Bagging Technique to Reduce Misclassification in Coronary Heart Disease Prediction Based on Random Forest. In Proceedings of the Journal of Physics: Conference Series; Institute of Physics Publishing: Bristol, UK, 2020; Volume 1477. [Google Scholar] [CrossRef] [Scilit]
- S., M.E.; Fajar, M.; T., M.I.; Jatmiko, W. FNGLVQ FPGA Design for Sleep Stages Classification based on Electrocardiogram Signal. In Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics; IEEE: New York, NY, USA, 2012. [Google Scholar]
- van Veen, R.; Gurvits, V.; Kogan, R.V.; Meles, S.K.; de Vries, G.J.; Renken, R.J.; Rodriguez-Oroz, M.C.; Rodriguez-Rojas, R.; Arnaldi, D.; Raffa, S.; et al. An application of generalized matrix learning vector quantization in neuroimaging. Comput. Methods Programs Biomed. 2020, 197, 105708. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fan, J.; Yue, W.; Wu, L.; Zhang, F.; Cai, H.; Wang, X.; Lu, X.; Xiang, Y. Evaluation of SVM, ELM and four tree-based ensemble models for predicting daily reference evapotranspiration using limited meteorological data in different climates of China. Agric. For. Meteorol. 2018, 263, 225–241. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Luan, F.; Wu, Y. A comparative assessment of six machine learning models for prediction of bending force in hot strip rolling process. Metals 2020, 10, 685. [Google Scholar] [CrossRef] [Scilit]
- Bargel, H.J.; Schulze, G. Werkstoffkunde, 11th ed.; Springer: Berlin/Heidelberg, Germany, 2012. [Google Scholar] [CrossRef] [Scilit]
- ABVML Project—Smart Materials Research. Available online: https://www.inw.hs-mittweida.de/webs/wfq/smart-materials/forschung-smart/abvml/ (accessed on 9 April 2025).







| Number of Samples | p Class (%) | n.p. Class (%) |
|---|---|---|
| 3599 | 93.5 | 6.5 |
| Type of Feature | Feature | Samples with This Feature (%) |
|---|---|---|
| Numerical | material thickness (mm) | 100.0 |
| Z (%) | 40.9 | |
| std. dev. of Z (%) | 40.8 | |
| CVN (J) | 19.8 | |
| std. dev. of CVN (J) | 19.8 | |
| Sn [wt%] | 4.0 | |
| B [wt%] | 8.6 | |
| C [wt%] | 18.9 | |
| Mn [wt%] | 18.9 | |
| Si [wt%] | 18.9 | |
| S [wt%] | 18.7 | |
| P [wt%] | 18.9 | |
| Cr [wt%] | 18.8 | |
| Ni [wt%] | 18.6 | |
| Cu [wt%] | 18.6 | |
| As [wt%] | 4.1 | |
| Ti [wt%] | 17.1 | |
| N [wt%] | 18.2 | |
| Al [wt%] | 18.7 | |
| Mo [wt%] | 15.6 | |
| V [wt%] | 17.0 | |
| Nb [wt%] | 17.5 | |
| Categorical | steel grade/material | 100.0 |
| delivery condition | 99.7 | |
| rolling direction | 100.0 |
| Model | Architecture | Remarks | Successful Applications |
|---|---|---|---|
| DT [13,14] | Tree structure via recursive feature partitioning; splits by metrics (e.g., information gain); stop by depth/size; pruning prevents overfitting | Good interpretability; risk of overfitting; flexible (classification and regression) | Fault cause analysis in manufacturing [22] |
| KNN [15] | Instance-based, non-parametric; prediction by majority vote of the k nearest neighbours; lazy learning | Suitable for multiclass classification; long computation times for large datasets | Diagnosis of wind turbine blade damage [23] |
| BCDT [16] | Ensemble of DTs; each tree trained on a bootstrap sample; final prediction by majority vote | Reduces variance and overfitting; improves stability and accuracy; insensitive to individual outliers | Plant identification [24] |
| RF [17] | Ensemble of DTs using bagging and random feature selection; final prediction by majority vote | Reduces variance and overfitting; delivers stable and robust predictions | Steel microstructure characterisation [2]; fault classification for induction motors [25]; cutting tool fault diagnosis [26] |
| BCRF [16] | Ensemble of RF base models; each RF trained on a bootstrap sample; final prediction by majority vote | Reduces variance, stabilises predictions, increases accuracy | Coronary heart disease prediction [27]; prediction of WBBT [12] |
| HGBC [18] | Gradient boosting: sequential weak learners (e.g., DTs); histogram-based binning for faster splits | Suitable for large, mixed datasets; high accuracy; efficient and fast training | Liver disease prediction [19] |
| GLVQ [20] | Prototype-based classification; classes represented by prototypes; assignment by relative distance | Robust performance; clear separation between class prototypes | Sleep stage classification from ECG [28] |
| GMLVQ [21] | GLVQ extension with learned distance metric; input assigned to nearest prototype | Improved performance; handles high-dimensional data; interpretable; weighs relevant features | Neurodegenerative disease classification [29] |
| Model | (%) | (%) | (%) |
|---|---|---|---|
| DT | 11.81 | 0.96 | 20.32 |
| RF | 20.15 | 2.48 | 32.40 |
| HGBC | 30.54 | 2.59 | 51.16 |
| KNN | 30.46 | 3.25 | 76.30 |
| BCDT | 13.81 | 23.18 | 2.23 |
| BCRF | 15.80 | 2.82 | 25.00 |
| GLVQ | 0.08 | 0.14 | 0.01 |
| GMLVQ | 2.24 | 1.28 | 3.17 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Backofen, F.; Hähnel, U.; Hahn, F.; Hockauf, K. Comparative Analysis of Various Supervised Machine Learning Models for the Prediction of the Outcome of the Welded Bead Bending Test. Metals 2026, 16, 418. https://doi.org/10.3390/met16040418
Backofen F, Hähnel U, Hahn F, Hockauf K. Comparative Analysis of Various Supervised Machine Learning Models for the Prediction of the Outcome of the Welded Bead Bending Test. Metals. 2026; 16(4):418. https://doi.org/10.3390/met16040418
Chicago/Turabian StyleBackofen, Fritz, Ulrike Hähnel, Frank Hahn, and Kristin Hockauf. 2026. "Comparative Analysis of Various Supervised Machine Learning Models for the Prediction of the Outcome of the Welded Bead Bending Test" Metals 16, no. 4: 418. https://doi.org/10.3390/met16040418
APA StyleBackofen, F., Hähnel, U., Hahn, F., & Hockauf, K. (2026). Comparative Analysis of Various Supervised Machine Learning Models for the Prediction of the Outcome of the Welded Bead Bending Test. Metals, 16(4), 418. https://doi.org/10.3390/met16040418

