Next Article in Journal
A Comprehensive Review of Metaheuristics for the Modern Traveling Salesman Problem and Drone-Assisted Delivery
Previous Article in Journal
MSA-Net: A Deep Learning Network with Multi-Axial Hadamard Attention and Pyramid Pooling for Stroke Microwave Imaging
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Robust Prediction of Compressive Strength of SCM Concrete with Nested Cross-Validation and Bayesian Optimization

1
Department of Architecture, Mimar Sinan Fine Arts University, 34427 Istanbul, Türkiye
2
Department of Civil Engineering, Istanbul University-Cerrahpasa, 34320 Istanbul, Türkiye
3
Department of Smart City, Gachon University, Seongnam 13120, Republic of Korea
*
Authors to whom correspondence should be addressed.
Algorithms 2026, 19(4), 277; https://doi.org/10.3390/a19040277
Submission received: 6 March 2026 / Revised: 27 March 2026 / Accepted: 30 March 2026 / Published: 2 April 2026

Abstract

Concrete production is one of the main sources of CO2 emissions. The primary reason for this is the high clinker content of Portland cement. To mitigate this problem, supplementary cementitious materials (SCMs) such as fly ash, silica fume, Ground granulated blast furnace slag (GGBFS), rice husk ash, and natural pozzolans are increasingly being used. These materials are used as partial replacements for cement. SCMs not only reduce the environmental impact of concrete but can also improve its long-term mechanical and durability properties. The aim of this study is to develop a machine learning framework that can accurately predict the compressive strength of concrete containing SCMs. The framework includes the training and evaluation of several machine learning models. Nested cross-validation and Bayesian hyperparameter optimization were used to explore the full capacity of the models and ensure reliable evaluation. Permutation significance testing and learning curve analysis were applied to verify that the models learn meaningful patterns rather than memorize the data. Also, feature importance and SHapley Additive exPlanations analyses were performed and the key variables that influence the prediction of the compressive strength of SCM concrete were identified. The optimized XGBoost model achieved the best generalization performance with a holdout R2 of 0.8398. It confirms the effectiveness of the proposed statistically rigorous machine learning framework for reliable compressive strength prediction of SCM-blended concrete.

1. Introduction

The concept of sustainability in the built environment revolves around conserving resources. It evaluates life cycle costs and designs for human comfort and usability. In terms of resource conservation, the 3Rs approach (Reduce, Reuse, and Recycle) is widely adopted within both construction and manufacturing industries [1]. As is known, concrete is a key building material due to its adaptability and design flexibility. Its embodied CO2 mainly depends on the cement content with concrete producing about 100 kg of CO2 per tonne on average. Since Portland cement is the main source of these emissions, reducing its use is the main essential key. Partial replacement with Supplementary Cementitious Materials (SCMs) such as Ground Granulated Blast Furnace Slag (GGBS), Pulverized Fly Ash (PFA), Rice Husk Ash (RHA), and Silica Fume (SF) can significantly lower the carbon footprint. SCMs are initially used to reduce costs, but now they are widely used to enhance the sustainability of concrete by decreasing CO2 emissions [2]. SCMs have long been integrated into concrete either as partial substitutes for clinker in cement or as replacements for a portion of cement in concrete mixtures. This approach benefits the industry by generally reducing costs, minimizing environmental impacts and enhancing both strength and durability.
Machine learning is particularly valuable for SCM-based concrete because the interactions between cement, supplementary materials, water content and curing conditions are highly nonlinear and difficult to represent using conventional empirical equations. By learning complex relationships directly from experimental datasets, machine learning models can provide more reliable and generalized predictions of compressive strength across diverse material compositions. Recently, several machine learning approaches have been proposed for SCMs. The related studies are mentioned as follows:
Sobuz et al. experimentally investigated the mechanical and durability properties of high-strength self-compacting concrete containing marble powder, fly ash, and silica fume, and used machine learning models to predict these properties. The results show that appropriate SCM ratios can increase the strength of concrete while reducing environmental impacts [3].
Moradi et al. investigated the mechanical effects of binary admixtures used in the concrete industry and developed a machine learning-based model to predict the compressive strength of concrete containing these materials. The proposed model can accurately predict the compressive strength of concrete with different SCM combinations and allows the determination of optimum binary SCM ratios [4].
Liu et al. modeled the creep behavior of additive-modified concrete using machine learning methods and analyzed the factors affecting this behavior. In the study, creep prediction models were developed using Linear Regression (LR), Support Vector Regression (SVR), Random Forest (RF), and Extreme Gradient Boosting (XGB) models. Parameter sensitivity was examined using SHapley Additive exPlanations (SHAP) analysis and the results showed that the XGB model could predict the creep fit of SCM concrete with high accuracy [5].
Ahmad et al. [6] comparatively investigated machine learning methods such as bagging, AdaBoost, gene expression programming and decision tree to predict the compressive strength of concrete containing fly ash and blast furnace slag. The results show that the bagging model can predict the compressive strength of concrete with high accuracy (R2 ≈ 0.92) and that ML methods can save time and resources in the concrete design process.
Miao et al. [7] used hybrid machine learning algorithms to predict the mechanical properties of concrete with recycled aggregate and admixtures. The results show that the SSA-XGB hybrid model can predict compressive, flexural and tensile strength and modulus of elasticity with high accuracy. Water/binder ratio, cement content, and superplasticizer amount are the most effective parameters.
Tran [8] comparatively investigated various machine learning algorithms to predict the chloride diffusion coefficient of concrete containing additives. The results show that the Gradient Boosting model provides the highest prediction performance. The basic parameters affecting chloride diffusion can be determined by using SHAP, Individual Conditional Expectation (ICE) and Partial Dependence Plot (PDP) analyses.
Kurniati et al. [9] developed machine learning models to predict the compressive strength of cement paste using copper mine waste as an admixture. The results show that the Random Forest model can predict compressive strength with high accuracy for different mix ratios and curing times, and that this approach can contribute to sustainable cement design.
Datta et al. [10] investigated the mechanical properties, microstructural behavior and environmental impacts of concrete using rice husk ash as an admixture using experimental and machine learning approaches. The results show that approximately 15% RHA substitution provides optimal strength and environmental performance. The XGBoost model offers high accuracy in predicting concrete properties.
Yoon et al. [11] used machine learning models to predict the reactivity of different admixtures depending on their physical and chemical properties. The results show that the artificial neural network (ANN) model provides the highest prediction accuracy. The important parameters can be determined in predicting SCM reactivity.
Golafshani et al. [12] proposed a framework integrating machine learning models with optimization algorithms to improve low-carbon mix design for concrete with recycled aggregates and admixtures. The results show that the CatBoost, XGBoost, and LightGBM models successfully predict the strength and durability properties of concrete. The proposed approach significantly reduces the carbon footprint.
Alkharisi and Dahish [13] used Response Surface Methodology (RSM) and machine learning models (M5P, Random Forest, XGBoost) together to predict the compressive strength of concrete with recycled aggregates and admixtures. The results show that the XGBoost model provides the highest prediction accuracy. Curing age and superplasticizer amount are the most influential parameters on compressive strength.
Chen et al. [14] comparatively investigated various machine learning models to predict the compressive strength of concrete with recycled aggregates containing fly ash and blast furnace slag. The results show that the particle swarm optimization-based backpropagation model provides the highest prediction accuracy, while curing age and cement content are the most important parameters determining strength development.
Sathiparan et al. [15] combined machine learning models with chemical composition data to predict the compressive strength of permeable concrete containing admixtures. The results show that the XGBoost model provides high prediction accuracy. CaO, SiO2 content, and curing time are critical parameters in strength prediction.
Parhi and Patro [16] combined machine learning with multi-objective optimization to predict and optimize the compressive strength, porosity, and cost of admixture concrete. The results show that the Random Forest model provides high accuracy in strength prediction (R2 = 0.97). the multilayer perceptron model provides high accuracy in porosity prediction (R2 = 0.96). The developed models allow for the determination of practical mix designs with Multi-Objective Differential Evolution Optimizer and a user-friendly graphical user interface (GUI).
Abbas and Muntean [17] used artificial neural networks (ANN) to predict the tensile strength of concrete containing waste-derived admixtures. The ANN model, which is optimized via the Levenberg–Marquardt algorithm, provided predictions with high accuracy (R = 1) and low MSE over 440 data points. The study shows that while SCMs improve concrete durability and environmental performance, the ANN-based approach can reduce the need for physical testing and accelerate sustainable mix design.
Although many current studies have successfully applied machine learning to predict the properties of concrete containing SCM, the complex nonlinear interactions between mix parameters and the need for reliable generalization highlight the importance of developing robust and statistically reliable machine learning frameworks. The key innovation of this study is the development of a statistically robust and fully integrated machine learning framework for predicting the compressive strength of SCM-blended concrete. This framework also aims to provide reliable generalization and interpretability.
Unlike many previous studies, this research does not rely solely on traditional cross-validation or simple training-test data separation. Instead, it employs a nested cross-validation architecture combined with Bayesian hyperparameter optimization to provide a clear separation between model tuning and performance evaluation. Furthermore, permutation-based significance tests are applied to verify that model performance stems from genuine input–output relationships, not random correlations or overfitting. In addition, a SHAP-based explainable artificial intelligence approach is used to transparently interpret the effect of mix components on strength prediction. The developed framework was validated using a large global dataset containing 1456 SCM concrete mixes. The results show that the proposed method is robust and reliable for data-driven strength prediction.
Recent studies have also increasingly focused on combining predictive accuracy with interpretability in machine learning-based modeling. For instance, interpretable surrogate models based on Random Forest algorithms, combined with SHAP analysis, have been proposed to explain model predictions while maintaining high accuracy [18].
In addition, several recent studies have explored advanced data-driven and hybrid learning frameworks, including multi-source data-driven models, hybrid machine learning approaches, and deep learning architectures such as CNN-LSTM optimized using Bayesian techniques [19,20,21]. These developments highlight a growing trend toward more robust, interpretable, and data-efficient predictive systems.
Compared to these recent advances, the present study integrates nested cross-validation, Bayesian hyperparameter optimization, permutation significance testing, and SHAP-based interpretability within a unified framework. This combination provides not only strong predictive performance but also improved statistical reliability and model transparency.

2. Materials and Methods

A large-scale global dataset was curated from the study “A Global Dataset of SCM-Blended Concrete and Its Properties” by Liu et al. [22] that was available on Mendeley Data. This dataset encompasses variables such as cement, water, fly ash, coarse aggregate, silica fume, ground granulated blast furnace slag (GGBFS) and superplasticizer for concrete mix proportion [23].
The methodology starts with exploratory data analysis, following this with a set of machine learning models that were evaluated using nested cross-validation and hyperparameter optimization. Then, a feature importance analysis and an SHAP analysis were performed on the best-performing models. Finally, the study concludes with permutation significance testing and learning curve analysis proving the models’ learning as genuine but not memorization of data.

2.1. The Dataset

2.1.1. Exploratory Data Analysis

The dataset size is 1456 × 9. This dataset consists of 8 independent variables and 1 target variable (Cylinder Compressive Strength). The input variables are cement, fly ash, water, superplasticizer, coarse aggregate, fine aggregate, silica fume, and ground granulated blast furnace slag. Figure 1 shows the distribution of each variable with histogram graphs. In addition to the histograms, kernel density estimation (KDE) curves are also provided. Descriptive statistics are also presented in each histogram. These are minimum and maximum values, mean (µ), standard deviation (σ), skewness (β), and kurtosis (γ).

2.1.2. Correlation Analysis

The correlation matrix shown in Figure 2 summarizes the linear relationships between concrete mix parameters and the obtained compressive strength. The correlation values shown in the matrix were derived from the dataset as Pearson correlations between features. This matrix, showing pairwise correlation coefficients, reveals how strongly each input variable relates to other variables. Therefore, it is a useful tool for identifying multicollinearity, understanding parameter interactions, and guiding feature selection for modeling. On this scale, values close to +1 indicate strong positive correlation. Values close to −1 indicate strong negative correlation. Coefficients close to zero indicate weak or no linear relationships.
As seen in Figure 2, there are clear reciprocal relationships between some mix components. For example, there is a clear positive correlation between cement and compressive strength. The water variable, however, shows a weak negative trend with both aggregates and strength. Similarly, coarse and fine aggregates show a moderate negative correlation with some binder materials. This indicates opposing interaction trends within the mix design.

2.2. Machine Learning Methods

In this study, machine learning (ML) techniques were applied using a comprehensive global dataset to model and predict the mechanical properties of concretes containing SCMs. Algorithms were trained and tuned on mixtures with different binder types and varying cementitious component ratios.
Each model underwent a careful hyperparameter optimization process to reduce prediction errors and improve generalization performance. Comparing the outputs of the models provides important insights into the ability of ensemble learning approaches to capture complex interactions between material components and accurately predict compressive strength. A brief introduction to each model used in the training and evaluation process is given below.

2.2.1. K Nearest Neighbors

One of the employed algorithms, K-Nearest Neighbors (KNN), is a non-parametric and instance-based learning approach that can be applied to both classification and regression problems. In this method, the value of a new observation is determined based on the closest data points in the training set using a selected distance metric. Although various distance measures can be used, Euclidean distance is the most common [24]. For regression tasks, KNN estimates the output by averaging the values of the k nearest neighbors, and the choice of k significantly influences model performance.
In this study, the KNN model was implemented using the Scikit-learn library within the proposed nested cross-validation framework. The mathematical formulation of the KNN method is given below:
y ^ = 1 k i = 1 k   y i ,

2.2.2. Multilayer Perceptron

Artificial Neural Networks (ANNs) have been widely used since the 1990s. Their prediction performance is influenced by several factors, including the selected input variables, network architecture, activation functions, and training algorithm [25]. Among different ANN structures, the Multilayer Perceptron (MLP) is commonly used due to its flexibility and wide range of applications [26]. However, the Levenberg–Marquardt (LM) algorithm, often used for training MLP models, may face limitations when applied to highly complex datasets [27]. More details on MLP structures can be found in Taud and Mas [28].
The MLP model was implemented within the proposed framework. A schematic representation of the MLP structure is shown in Figure 3.

2.2.3. Support Vector Machine

Another employed algorithm, Support Vector Machine (SVM), is regarded as one of the key algorithms in machine learning. Its strength lies in its capacity to identify classification patterns with reliable accuracy and consistent performance. While it can also be applied to certain regression tasks, SVM is predominantly recognized today as a robust and widely adopted method for classification problems [29].

2.2.4. ElasticNet

The used algorithm called Elastic Net Regression [30] is a regularization technique that combines Lasso (L1) and Ridge (L2) penalties to address multicollinearity and reduce overfitting in high-dimensional datasets. Its mathematical formulation is expressed as follows:
y = b0 + b1 × x1 + b2 × x2 + … + bn × xn + e,

2.2.5. Decision Tree

Decision Trees (DTs) [31] were also employed in this study. These models are hierarchical structures widely used for classification and regression tasks, where decisions are made sequentially based on input features. Compared to neural networks, DTs offer a simpler and more interpretable structure.
The Decision Tree model was implemented within the proposed framework. A schematic representation of the DT structure is shown in Figure 4.

2.2.6. Bagging

Bagging (bootstrap aggregation) is an ensemble learning technique that generates multiple versions of a base model by training on different bootstrap samples and combining their outputs. For regression tasks, predictions are averaged, while for classification, a majority voting scheme is used. This approach is particularly effective for reducing variance and improving the stability of models that are sensitive to variations in the training data [32]. In this study, the Bagging approach was implemented using Decision Trees as base learners.

2.2.7. Random Forest

The Random Forest (RF) algorithm is a widely used ensemble method for both classification and regression tasks [33]. It constructs multiple decision trees using randomly sampled subsets of the data and combines their outputs through averaging or majority voting [34,35]. RF is known for its robustness against overfitting and its ability to handle high-dimensional datasets, although it may require careful hyperparameter tuning [36].
In this study, the Random Forest model was implemented within the proposed framework. A schematic representation of the RF structure is shown in Figure 5. The node colors indicate variations in the predicted continuous values in the regression trees.

2.2.8. Extreme Gradient Boosting (XGBoost)

XGBoost is an ensemble learning algorithm based on decision trees, where multiple trees are constructed sequentially, and each tree is trained to correct the errors of the previous ones, resulting in high predictive accuracy [37]. This employed algorithm was chosen for its proficiency in managing tabular datasets and its capacity to model complex interactions among features without requiring a predefined order.
The iterative mechanism of XGBoost can be expressed as in Equation (3), where fk represents the weak learner at the kth iteration, and N denotes the total number of weak learners.
f ( x ) = K = 1 N   f k ( x ) = y ^ ,
During each iteration, the model seeks to minimize the objective function J presented in Equation (4), where M is the total number of training samples, l is the loss function, yi and y ^ i are the observed and predicted values and Ω is a regularization term. Figure 6 shows the stepwise process of XGBoost, highlighting the residuals εk at iteration k.
J = i = 1 M 1 ( y i , y ^ i ) + K = 1 N   ( f k )

2.2.9. Categorical Boosting (CatBoost)

CatBoost was also used in this study. It uses greedy algorithms to improve prediction accuracy. In this method, features are ranked according to the way they create splits in the trees and are then applied within each leaf. The depth of the tree and other structural limits are set in advance through the model’s parameters. Before each new tree is built, the data used for classification or regression is randomly permuted.
CatBoost evaluates the model at every step with a metric that guides how the function should be improved when constructing the next tree. In many standard tests, CatBoost has outperformed well-known gradient boosting packages and has achieved strong results on widely used benchmarks [38,39].

2.2.10. Light Gradient Boosting Machine (LightGBM)

LightGBM is a gradient boosting decision tree (GBDT) algorithm introduced in 2017. It applies to both classification and regression tasks [40]. Compared to traditional GBDT, LightGBM trains models much faster by employing several optimizations. It uses a histogram-based approach that bins continuous features into discrete intervals, significantly reducing computation time for tree construction, especially with large datasets [41].

2.3. Hyperparameter Optimization with Optuna

In ensemble-based learning algorithms, model performance largely depends on hyperparameter selection. Instead of default or manually selected values, a systematic hyperparameter optimization has been implemented to ensure more stable operation of the models and a fair comparison between algorithms.
In this study, hyperparameter tuning was performed using the Optuna optimization framework, which is based on the sequential model-based optimization approach [42]. Optuna uses the Tree-structured Parzen Estimator (TPE) algorithm to efficiently navigate complex and high-dimensional search spaces [43].
Unlike grid search methods that examine all predefined combinations or random search approaches that perform random sampling [44], the TPE-based method dynamically updates the sampling distribution based on previous results. Thus, the search process increasingly focuses on promising regions in the parameter space, reducing unnecessary computational costs.
Hyperparameter optimization was performed only on the training dataset. This choice eliminates the risk of information leakage and ensures that the test data remains completely unseen during the model development process. A separate search space was defined for each algorithm. The most important parameters included the number of estimators, maximum tree depth, learning rate, minimum number of samples required for splitting and leaf nodes, subsampling rates, and regularization coefficients.
Each trial configuration was evaluated using k-fold cross-validation within the training subset. The optimization goal was to minimize the selected error metric (RMSE or MAE according to the evaluation criteria used in this study). The final hyperparameter set for each model was selected based on the lowest mean cross-validation error.
A fixed random seed (42) was used throughout the optimization process to increase reproducibility. Furthermore, overfitting was reduced and generalization stability was increased by applying early stopping criteria where supported by the algorithm.
In the implementation the sampler was Tree-structured Parzen Estimator (TPE) of optuna, utilized via optuna.samplers.TPESampler (seed = 42), the optimization direction was maximizing R2, trials per outer fold was 100, pruner was none (all trials run to completion) and the objective metric: Mean 5-fold cross-validated R2 on the outer training split. The range of search spaces was defined by using Optuna’s trial.suggest_* API. Table 1 below provides information on the search spaces for all the optimized hyperparameters.
All models were wrapped in a sklearn.pipeline. Pipeline with StandardScaler as the first step (applied inside each CV fold to prevent data leakage). Hyperparameter selection uses the best trial from the last outer fold, and that final models are refit on the full training set with those parameters before holdout evaluation.
The optimized parameters were used to train the final models, and the models were evaluated on an independent test dataset. This structured and data-driven tuning strategy has improved model stability and generalization performance while providing transparent and reproducible comparisons between methods.

2.4. Nested Cross Validation

A nested cross-validation (CV) framework was used to obtain an unbiased estimate of model performance during hyperparameter optimization. First, the dataset was divided into two parts: an 80% training dataset and a 20% independent holdout dataset. The holdout dataset remained untouched until the final evaluation stage.
Then, the training dataset underwent a 10-fold cross-validation process in the outer loop. In each iteration, one fold was reserved as an outer validation subset, while the remaining folds formed the outer training subset. Hyperparameter optimization was performed within each outer training subset using Optuna TPE with 100 trials. This process was carried out with 5-fold cross-validation in the inner loop.
In this inner loop, candidate hyperparameter configurations were evaluated only on the outer training data. Thus, the outer validation fold remained completely unseen throughout the tuning process. After determining the best hyperparameters at the end of the inner loop, the model was retrained using the entire outer training subset and evaluated on the relevant outer validation fold. This process generated an unbiased out-of-fold (OOF) performance estimate.
This process was repeated for all outer folds, and the resulting validation scores were averaged to calculate the final cross-validation performance.
After completing the nested CV, the model was retrained on the entire training dataset using the hyperparameters obtained from the last outer fold. Since hyperparameters can vary between outer folds, the configuration obtained from the last outer fold was chosen. This approach is a common method in nested CV applications. The final model was then evaluated once on an independent holdout dataset to obtain the final generalization performance. The nested CV structure is provided in Table 2.
Within each outer fold, 100 Optuna trials were conducted per algorithm, each evaluated using 5-fold cross-validated R2 on the outer training subset, resulting in 1000 evaluations per algorithm. The final model was refitted on the full training dataset using a representative hyperparameter configuration derived from the nested CV process and evaluated on the independent holdout set.
This nested CV design provides a clear separation between hyperparameter optimization and performance evaluation. Thus, optimistic bias is avoided, and the model’s true predictive ability is reliably measured. All models were implemented in Python using python libraries. Library names and versions used in the implementation were as follows: Python (3.13.12), numpy (2.2.6), pandas (2.3.3), matplotlib (3.10.6), seaborn (0.13.2), scikit-learn (1.7.2), scipy (1.16.2), xgboost (3.0.5), lightgbm (4.6.0), catboost (1.2.8), optuna (4.5.0), shap (0.50.0).
A description of the nested CV metrics is given in Table 3.
In order to facilitate reproducibility a single global constant random_state = 42 was used throughout the entire code, applied constantly to:
  • The 80/20 stratified train/holdout split (train_test_split, random_state = 42);
  • The outer 10-fold CV splitter (KFold, shuffle = True, random_state = 42);
  • The inner 5-fold CV splitter (KFold, shuffle = True, random_state = 42);
  • All nine model instantiations that accept a random state argument;
  • The Optuna TPE sampler (TPESampler (seed = 42));
  • The SHAP background data sampler (np.random.default_rng (42));
  • The permutation test label shuffler (np.random.default_rng (42));
  • The learning curve CV splitter (KFold, shuffle = True, random_state = 42).

2.5. Shapley Additive Explanations (SHAP) Method

The SHAP (Shapley Additive Explanations) method is based on cooperative game theory and the concept of Shapley value. In the context of machine learning, input variables can be thought of as “players” in a cooperative game. The prediction model represents the “game” to be solved.
In this approach, the contribution of each variable to a single prediction is determined by calculating its average marginal effect on all possible combinations of variables. Because all possible feature combinations are considered, SHAP theoretically distributes variable importance in a robust and fair manner. Thus, the contribution assigned to each variable is both consistent and mathematically meaningful.
The SHAP framework [45] provides important theoretical properties such as local accuracy, consistency, and missingness. The local accuracy property guarantees that the sum of all variable contributions is exactly equal to the model output for a given sample. The consistency property guarantees that if the effect of a variable in a model increases, the importance value assigned to that variable will not decrease.
These properties make SHAP particularly suitable for explaining complex nonlinear models that are difficult to interpret. It offers a powerful explanation method, especially for ensemble-based tree algorithms.
In this study, SHAP analysis was applied to trained regression models. The aim is to better understand how mix design parameters affect compressive strength predictions. Both global interpretations based on mean absolute SHAP values and local explanations for individual samples have been examined.
This two-level analysis allowed for a detailed assessment of the relative effects of binder composition, SCMs and other input variables. Integrating the SHAP method into the model facilitated the engineering interpretation of model behavior and increased the reliability of the proposed prediction approach.

2.6. Permutation Significance Test

To verify that model performance reflects genuine input–output relationships rather than coincidental overfitting or data leakage, a permutation significance test was conducted for all nine regression algorithms. The target variable (compressive strength, MPa) was randomly shuffled 100 times while the feature matrix was held fixed; each model was retrained from scratch under identical hyperparameter settings and evaluated on the original holdout set. The explanation of Permutation Significance Test Metrics is shown in Table 4.
In this procedure, the target variable was randomly shuffled (N) times while keeping the feature matrix fixed. For each permutation i 1 , , N the model was retrained from scratch using the same hyperparameters and evaluated on the unchanged holdout dataset, yielding a performance score T i perm . The performance of the original model trained on the true (unshuffled) data is denoted as T real .
The permutation test p-value is then computed as:
p = 1 + i = 1 N I T i perm T real N + 1
where I is the indicator function that returns 1 if the condition is satisfied and 0 otherwise. This formulation estimates the probability of obtaining a performance equal to or greater than the observed model performance under the null hypothesis that no relationship exists between the input features and the target variable.
For regression metrics where higher values indicate better performance (i.e., R2), the comparison T i perm T real is used. Conversely, for error-based metrics (e.g., RMSE or MAE), the inequality is reversed. A p-value below a predefined significance threshold (p < 0.05) indicates that the model captures statistically significant structure in the data and that its performance is unlikely to have occurred by random chance.

2.7. Learning Curve Analysis

To assess whether each model genuinely learns from data or merely memorizes training patterns, learning curves were constructed by progressively increasing the training set size from 10% to 100% of the available training data (in five steps: 10%, 30%, 50%, 70%, and 100%), with train and validation R2 recorded at each step via 10-fold cross-validation. A well-generalizing model exhibits a characteristic pattern: training R2 decreases (i.e., can be a slight decrease) as more samples are added (reduced memorization), while validation R2 increases and the two curves converge toward a common plateau, indicating that additional data reduces the bias–variance gap. In contrast, a high-variance (overfitting) model maintains near-perfect training R2 across all sizes (with no decrease) while validation R2 plateaus at a significantly lower level, with a persistent gap between the two curves.

3. Results

3.1. Model Training with Nested CV and Bayesian Hyperparameter Optimization

In the training stage, as explained earlier, the best (optimized) hyperparameters were derived from last outer fold. Table 5 below presents the best hyperparameters identified for each model.
The results of the model training and evaluation process with nested cross-validation and Bayesian hyperparameter optimization are provided in Table 6. The CV Train, and CV Val metrics were obtained from the outer folds as mean values of all folds of Nested CV. The CV Val scores represent the results from data not seen for that fold by the hyperparameter optimizer, which is running on the inner folds. The holdout metrics were obtained from the holdout subset never seen during the training process.
Within the nested cross-validation framework employed in this study, the outer training and outer validation performances serve fundamentally different roles in assessing model behavior. The outer training performance reflects the model’s ability to fit the data used for parameter estimation after hyperparameter optimization has already been completed in the inner loop. In contrast, the outer validation performance represents a strictly unbiased estimate of predictive capability, as these validation folds remained completely unseen during both model training and hyperparameter tuning. Because hyperparameters were optimized exclusively using the inner cross-validation procedure applied to the outer training subset, the outer validation folds were fully isolated from the optimization process. Consequently, the performance measured on the outer validation folds provides a realistic simulation of model behavior on new, unseen data during the training process. The higher performance observed on the outer training subsets compared to the outer validation subsets is therefore expected and reflects the natural consequence of model fitting, rather than methodological bias or information leakage.
In the present results, models such as XGBoost, CatBoost, and LightGBM achieved very high outer training R2 values, approaching 0.99 due to models’ capacity optimization with alignment to the data in that fold with hyperparameter tuning, while their outer validation R2 values remained consistently high, ranging approximately between 0.86 and 0.88. This difference arises because the outer training subsets are directly used to estimate model parameters, allowing the model to more closely align with the specific structure of those data. In contrast, the outer validation subsets represent entirely independent samples that were not involved in either hyperparameter selection or parameter estimation.
The fact that validation performance remained consistently high despite this strict independence demonstrates that the models captured generalizable relationships rather than dataset-specific or fold-specific artifacts. A similar pattern was observed for ensemble models such as Random Forest and Bagging, which exhibited strong training performance accompanied by only moderate reductions in outer validation performance. This behavior confirms the effectiveness of ensemble learning in balancing model flexibility and generalization.
Conversely, simpler models such as Elastic Net showed smaller differences between training and validation performance, reflecting their more constrained functional form and lower capacity to fit complex patterns. Importantly, the nested cross-validation design used in this study eliminates optimistic bias because hyperparameter tuning, model fitting, and performance evaluation were performed on strictly separated data partitions. Therefore, the observed differences between outer training and outer validation performance represent the true and expected generalization gap inherent to supervised learning, rather than an artifact of improper evaluation.
Figure 7 and Figure 8 show the predicted vs. true diagrams for validation holds and holdout subset. A critical indicator of model generalization is the degree of agreement between cross-validation performance and independent holdout performance.
In this study, the comparison of CV Validation R2 and Holdout R2 values across all models revealed a consistent and interpretable pattern. Among the gradient boosting models, XGBoost achieved a CV Validation R2 of 0.8666 and a Holdout R2 of 0.8398. It has a small difference of 0.0268. Similarly, CatBoost produced a CV Validation R2 of 0.8792 and a Holdout R2 of 0.8367 with a difference of 0.0425.
LightGBM showed a CV Validation R2 of 0.8651 and Holdout R2 of 0.8239. This yields a difference of 0.0412. These relatively small reductions indicate that the validation performance estimates were realistic and did not substantially overestimate model performance on unseen data.
The ensemble tree-based models demonstrated comparable behavior. Random Forest yielded CV Validation R2 of 0.8330 and Holdout R2 of 0.8121 (difference: 0.0209), while Bagging produced CV Validation R2 of 0.8125 and Holdout R2 of 0.7975 (difference: 0.0150). These minimal differences suggest highly stable generalization performance and confirm the robustness of ensemble methods in this dataset.
The Decision Tree model showed nearly identical performance between validation and holdout datasets with CV Validation R2 of 0.7146 and Holdout R2 of 0.7156, indicating virtually perfect agreement. This reflects low variance behavior but also confirms its comparatively limited predictive capacity relative to ensemble methods.
The neural network model (MLP) demonstrated moderate consistency with CV Validation R2 of 0.8019 and Holdout R2 of 0.7791 (difference: 0.0228). It is an acceptable generalization without substantial overfitting.
The K-Nearest Neighbors model exhibited the largest reduction, with CV Validation R2 of 0.7894 and Holdout R2 of 0.7523 (difference: 0.0371), suggesting relatively higher sensitivity to data variation and a slightly optimistic cross-validation estimate. Finally, Elastic Net regression showed an unusual but favorable pattern, where Holdout R2 (0.5977) slightly exceeded CV Validation R2 (0.5738). This indicates that cross-validation provided a conservative estimate of model performance rather than an optimistic one.
Overall, the differences between CV Validation R2 and Holdout R2 remained consistently small across all models, generally within the range of 0.01 to 0.04. Such close agreement confirms that the cross-validation procedure provided reliable and unbiased performance estimates. More importantly, the absence of large performance drops between validation and holdout datasets indicates that the models successfully learned generalizable relationships rather than memorizing the training data. Among all methods, gradient boosting and ensemble tree models demonstrated the best balance between predictive accuracy and generalization stability, confirming their suitability for this regression problem.
The CatBoost appeared as the best model based on the CV Validation R2/RMSE, but based on the Holdout R2/RMSE the XGBoost exhibited a slightly better performance. As the holdout performance is on an unseen data subset, we consider fine-tuned XGBoost as the best-performing model for this dataset.
Figure 9 illustrates the CV Validation R2 for each Outer Validation Fold of Train Subset. In folds 2 and especially 9, the models struggle to reach high metrics, and models exhibit a slight drop of performance in fold 5. The sharp decrease observed in fold 9 may be attributed to the presence of more complex validation samples or potential outliers within that fold. Other fold scores follow a flat line pattern, indicating that the model generalizes consistently in the remaining 7 folds.

3.2. Feature Importances and Model Interpretation

The comparative analysis of feature importance across XGBoost, CatBoost, LightGBM, and Random Forest models (Figure 10) revealed a highly consistent pattern in identifying the dominant factors governing the compressive strength of concrete containing SCMs, while also highlighting subtle model-specific differences in feature prioritization. Across three of four models, cement content and water content were consistently identified as the two most influential predictors, confirming their fundamental role in controlling concrete strength. In particular, Random Forest assigned the highest relative importance to the concrete mix proportion of cement (0.293), followed by water (0.204), while CatBoost similarly emphasized water (0.241) and cement (0.230) as the primary governing variables. XGBoost also demonstrated comparable trends, with cement and water ranked among the top predictors. This strong consensus across different algorithmic architectures confirms the dominant influence of binder quantity and water-to-binder ratio on compressive strength development.
The feature importance rankings obtained using the XGBoost internal importance metric and SHAP analysis (Figure 11) showed some minor differences due to their fundamentally different computational principles. The XGBoost importance metric reflects the contribution of each feature to reducing the loss function in the tree-building process, while the SHAP importance represents the average magnitude of each feature’s contribution to the model’s prediction output. Therefore, some features, such as GGBFS, showed relatively lower importance in the tree-based importance metric but higher importance in the SHAP analysis. This indicates that GGBFS, although used less frequently in tree splitting, has a significant impact on the prediction results when used. In contrast, features such as SP and fine aggregate, although used more frequently in tree splitting, have a relatively lower overall impact on the prediction magnitude. These findings demonstrate that the SHAP method interprets feature impact more comprehensively and reliably because it directly quantifies the contribution of each feature to the model predictions rather than relying solely on the tree structure [46].
The SHAP summary analysis revealed that the most important parameter affecting compressive strength estimation was cement content, with higher cement amounts positively contributing to strength improvement. In contrast, water content showed an inverse relationship, with increasing water content negatively impacting the predicted strength. Complementary cementitious materials, particularly silica fume (SF) and GGBFS, showed positive effects at high dosages, confirming their beneficial role in increasing concrete strength. Aggregate-related variables and superplasticizer (SP) showed moderate effects, while the effect of fly ash (FA) remained relatively limited. This observation may be related to the characteristics of the dataset, since the influence of fly ash on compressive strength strongly depends on curing conditions and age. In particular, insufficient curing of fly ash-based concrete can significantly reduce compressive strength due to the slower pozzolanic reaction of FA compared to Portland cement. Higher aggregate amounts were generally associated with slightly negative or neutral SHAP values. High aggregate contents resulted in a decrease in estimated compressive strength, reflecting a dilution of the binder content in the mix. On the other hand, high SP content was more frequently associated with positive contributions, demonstrating its role in indirectly increasing strength by improving workability and particle packing. While high FA values sometimes provided positive contributions, their overall effect remained limited compared to other complementary cementitious materials. Overall, the SHAP analysis confirms that the model successfully captures the fundamental physical mechanisms determining concrete strength.

3.3. Permutation Significance Test

The permutation test results (Table 7) confirmed that all machine learning models learned statistically significant relationships and did not simply memorize noise. Model performance decreased significantly when the target variable was replaced with a random permutation, with R2 values dropping to negative values in the permuted models, compared to approximately 0.84 in the original models. This significant performance decrease indicates that the models’ predictive ability is entirely based on the actual relationships between the input variables and compressive strength. Furthermore, the permutation test produced a p-value of 0.0 for all models, confirming that the observed performance is statistically significant and not dependent on random chance. These findings provide strong evidence that the models perform genuine learning without overfitting.

3.4. Learning Curve Analysis

The learning curve analysis of all machine learning models (Figure 12) revealed consistent and robust generalization behavior. Validation performance steadily increased as the training set size increased, approaching a plateau level at larger sample sizes. Boosting models like XGBoost, CatBoost, and LightGBM have demonstrated near-perfect training and high validation performance, with the difference between training and validation curves decreasing as the dataset increases, showing strong learning capacity without severe overfitting. Similarly, ensemble methods like Random Forest and Bagging have shown learning behavior where validation performance increases steadily, and a moderate, controlled difference is maintained relative to training performance. This confirms a well-balanced bias–variance trade-off.
Decision Tree and KNN models started with a larger training–validation difference at smaller sample sizes, reflecting high variance and sensitivity to limited data; however, this difference decreased significantly as the dataset increased, and generalization ability improved. The MLP model showed a significant improvement in both training and validation performance as the dataset size increased, demonstrating that sufficient data is necessary for the neural network to effectively learn fundamental relationships. In contrast, the ElasticNet model showed poor performance with training and validation curves almost overlapping, experiencing underfitting due to limited model complexity, not overfitting. It is important to note that the convergence of the validation curves in all models and the decrease in performance variance over larger training sets indicate that the dataset is large enough to capture the underlying predictive relationships.
Overall, these findings confirm that the advanced ensemble and boosting models achieved the best generalization performance while maintaining robustness and avoiding overfitting, whereas simpler models were primarily limited by underfitting rather than excessive variance.

4. Conclusions

We provided a robust machine learning-based framework for estimating the compressive strength of SCM-blended concrete on a comprehensive global dataset of concrete mixtures containing cement-based materials. Considering that the cement industry contributes greatly to CO2 emissions, increasing the use of SCM such as fly ash, silica fume, high furnace slag, and natural light stone is not only an environmental necessity but also a means to improve the long-term properties of concrete.
In this study, a new pipeline including a nested cross-validation framework combined with Bayesian hyperparameter optimization was implemented to ensure the optimal performance of each model can be determined and evaluated for this dataset. Furthermore, the framework included a permutation significance testing and Learning Curve Analysis to ensure that the models evaluated learned genuine relationships rather than random patterns. The comparison between cross-validation and independent holdout performance, together with learning curve analysis, was used to confirm the absence of overfitting and to demonstrate that the models achieved true generalization. Finally, SHAP-based explainable artificial intelligence was utilized to interpret the influence of individual mixture components on strength prediction.
The most highly performing models include XGBoost, which achieved a CV validation R2 of 0.8666 and a holdout R2 of 0.8398, followed by CatBoost and LightGBM with very similar performance metrics. The small differences observed between XGBoost and CatBoost indicate that both models demonstrated similarly strong predictive performance, while XGBoost showed a slight advantage (0.8398 vs. 0.8367 R2) on the holdout dataset, this difference may depend on the specific data split and does not necessarily indicate that there exists a universally superior model.
The results suggest that ensemble machine learning models are highly effective in capturing the underlying patterns in the dataset and accurately predicting the compressive strength of SCM concrete.
In addition, the SHAP analysis revealed that cement content was the most influential parameter, showing a positive relationship with compressive strength. Water content was identified as the second-most influential variable, exhibiting an inverse relationship. SCMs, particularly silica fume (SF) and GGBFS, showed positive contributions at higher dosages, confirming their beneficial role in enhancing concrete strength.
The main limitation of this study is that the models were developed using a global dataset containing only mixture composition variables, without incorporating external factors such as curing conditions, temperature, and humidity, which may influence concrete strength and limit generalizability under different field conditions. These factors will be tested with a new dataset in our feature research. In addition, future research will not only focus on the 28-day intensity, but also on the development of an interpretation model that can predict the intensity of the expression curve.

Author Contributions

Conceptualization, G.B. and F.A.; methodology, Ü.I.; software, Ü.I.; validation, Ü.I. and G.B.; investigation, S.M.N.; data curation, F.A.; writing—original draft preparation, S.M.N., F.A. and Ü.I.; writing—review and editing, S.M.N., G.B. and Z.W.G.; visualization, Ü.I. and F.A.; supervision, Z.W.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

Data is available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Bramley, G.; Power, S. Urban form and social sustainability: The role of density and housing type. Environ. Plan. B Plan. Des. 2009, 36, 30–48. [Google Scholar] [CrossRef]
  2. Concrete Industry Sustainability Performance Report, Surrey, UK. Available online: http://www.sustainableconcrete.org.uk/ (accessed on 15 February 2011).
  3. Sobuz, M.H.R.; Aditto, F.S.; Datta, S.D.; Kabbo, M.K.I.; Jabin, J.A.; Hasan, N.M.S.; Khan, M.M.H.; Rahman, S.A.; Raazi, M.; Zaman, A.A.U. High-strength self-compacting concrete production incorporating supplementary cementitious materials: Experimental evaluations and machine learning modelling. Int. J. Concr. Struct. Mater. 2024, 18, 67. [Google Scholar] [CrossRef]
  4. Moradi, N.; Tavana, M.H.; Habibi, M.R.; Amiri, M.; Moradi, M.J.; Farhangi, V. Predicting the compressive strength of concrete containing binary supplementary cementitious material using machine learning approach. Materials 2022, 15, 5336. [Google Scholar] [CrossRef]
  5. Liu, Y.; Li, Y.; Mu, J.; Li, H.; Shen, J. Modeling and analysis of creep in concrete containing supplementary cementitious materials based on machine learning. Constr. Build. Mater. 2023, 392, 131911. [Google Scholar] [CrossRef]
  6. Ahmad, W.; Ahmad, A.; Ostrowski, K.A.; Aslam, F.; Joyklad, P.; Zajdel, P. Application of advanced machine learning approaches to predict the compressive strength of concrete containing supplementary cementitious materials. Materials 2021, 14, 5762. [Google Scholar] [CrossRef]
  7. Miao, X.; Zhu, J.X.; Zhu, W.B.; Wang, Y.; Peng, L.; Dong, H.L.; Xu, L.Y. Intelligent prediction of comprehensive mechanical properties of recycled aggregate concrete with supplementary cementitious materials using hybrid machine learning algorithms. Case Stud. Constr. Mater. 2024, 21, e03708. [Google Scholar] [CrossRef]
  8. Tran, V.Q. Machine learning approach for investigating chloride diffusion coefficient of concrete containing supplementary cementitious materials. Constr. Build. Mater. 2022, 328, 127103. [Google Scholar] [CrossRef]
  9. Kurniati, E.O.; Zeng, H.; Latypov, M.I.; Kim, H.J. Machine learning for predicting compressive strength of sustainable cement paste incorporating copper mine tailings as supplementary cementitious materials. Case Stud. Constr. Mater. 2024, 21, e03373. [Google Scholar] [CrossRef]
  10. Datta, S.D.; Sarkar, M.M.; Rakhe, A.S.; Aditto, F.S.; Sobuz, M.H.R.; Shaurdho, N.M.N.; Nijum, N.J.; Das, S. Analysis of the characteristics and environmental benefits of rice husk ash as a supplementary cementitious material through experimental and machine learning approaches. Innov. Infrastruct. Solut. 2024, 9, 121. [Google Scholar] [CrossRef]
  11. Yoon, J.; Yonis, A.; Park, S.; Rajabipour, F.; Pyo, S. Prediction of the R 3 test-based reactivity of supplementary cementitious materials: A machine learning approach utilizing physical and chemical properties. Int. J. Concr. Struct. Mater. 2024, 18, 75. [Google Scholar] [CrossRef]
  12. Golafshani, E.M.; Behnood, A.; Kim, T.; Ngo, T.; Kashani, A. A framework for low-carbon mix design of recycled aggregate concrete with supplementary cementitious materials using machine learning and optimization algorithms. In Structures; Elsevier: Amsterdam, The Netherlands, 2024; Volume 61, p. 106143. [Google Scholar] [CrossRef]
  13. Alkharisi, M.K.; Dahish, H.A. The application of response surface methodology and machine learning for predicting the compressive strength of recycled aggregate concrete containing polypropylene fibers and supplementary cementitious materials. Sustainability 2025, 17, 2913. [Google Scholar] [CrossRef]
  14. Chen, X.; Chen, N.; Cheng, S.; Li, N.; Wu, Q.; Liu, Z.; Chen, B.; Zhang, Z. Machine learning models for predicting the compressive strength of sustainable recycled aggregate concrete incorporating supplementary cementitious materials. Sustain. Mater. Technol. 2026, 47, e01907. [Google Scholar] [CrossRef]
  15. Sathiparan, N.; Jeyananthan, P.; Subramaniam, D.N. A comparative study of machine learning techniques and data processing for predicting the compressive strength of pervious concrete with supplementary cementitious materials and chemical composition influence. Next Mater. 2025, 9, 100947. [Google Scholar] [CrossRef]
  16. Parhi, S.K.; Patro, S.K. Data-driven prediction and intelligent optimization of strength, porosity and cost of concrete with supplementary cementitious materials. J. Struct. Integr. Maint. 2025, 10, 2567084. [Google Scholar] [CrossRef]
  17. Abbas, M.M.; Muntean, R. Predictive modelling of concrete tensile strength using ANN and Supplementary Cementitious Materials (SCMs). Constr. Build. Mater. 2025, 489, 142362. [Google Scholar] [CrossRef]
  18. Chen, F.; Tan, B.; Tang, H.; Zhang, H.; Luo, Y.; Xiao, X.; Liu, Y.; Lu, N. An interpretable random forest surrogate for rapid SIF prediction and fatigue life assessment of double-sided U-rib welds in orthotropic steel decks. Eng. Fail. Anal. 2026, 187, 110582. [Google Scholar] [CrossRef]
  19. Zhang, H.; Zhao, L.; Chen, F.; Luo, Y.; Xiao, X.; Liu, Y.; Deng, Y. A Machine Learning and Multi-Source Authentic Data-Driven Framework for Accurate Fatigue Life Prediction of Welds in Existing Steel Bridge Decks. Thin-Walled Struct. 2026, 222, 114559. [Google Scholar] [CrossRef]
  20. Zhang, H.; Yang, X.; Luo, Y.; Chen, F.; Lu, N.; Liu, Y.; Deng, Y. A novel hybrid algorithm for damage detection in bridge foundations under complex underwater environments using ROV capture pictures. Eng. Struct. 2026, 352, 122131. [Google Scholar] [CrossRef]
  21. Xiao, X.; Tang, H.; Zhang, H.; Luo, Y.; Lu, N.; Liu, Y.; Chen, F. Bayesian optimization CNN-LSTM neural network for fatigue life prediction of Rib-to-deck Welds in orthotropic steel decks. In Structures; Elsevier: Amsterdam, The Netherlands, 2026; Volume 84, p. 111088. [Google Scholar] [CrossRef]
  22. Liu, W.; Mo, H.; Lee, C.K. A Global Dataset of SCM-Blended Concrete and Its Properties; Version 2; Mendeley Data: Amsterdam, The Netherlands, 2025. [Google Scholar] [CrossRef]
  23. Liu, W.; Mo, H.; Lee, C.K. Comprehensive dataset for performance evaluation of SCM-blended concrete. Data Brief 2025, 60, 111616. [Google Scholar] [CrossRef] [PubMed]
  24. Barjouei, H.S.; Ghorbani, H.; Mohamadian, N.; Wood, D.A.; Davoodi, S.; Moghadasi, J.; Saberi, H. Prediction performance advantages of deep machine learning algorithms for two-phase flow rates through wellhead chokes. J. Pet. Explor. Prod. 2021, 11, 1233–1261. [Google Scholar] [CrossRef]
  25. Misra, J.; Saha, I. Artificial neural networks in hardware: A survey of two decades of progress. Neurocomputing 2010, 74, 239–255. [Google Scholar] [CrossRef]
  26. Maier, H.R.; Dandy, G.C. Neural Networks for the Prediction and Forecasting of Water Resources Variables: A Review of Modelling Issues and Applications. Environ. Model. Softw. 2000, 15, 101–124. [Google Scholar] [CrossRef]
  27. Abiodun, O.I.; Jantan, A.; Omolara, A.E.; Dada, K.V.; Umar, A.M.; Linus, O.U.; Arshad, H.; Kazaure, A.A.; Gana, U.; Kiru, M.U. Comprehensive review of artificial neural network applications to pattern recognition. IEEE Access 2019, 7, 158820–158846. [Google Scholar] [CrossRef]
  28. Taud, H.; Mas, J.F. Multilayer perceptron (MLP). In Geomatic Approaches for Modeling Land Change Scenarios; Springer International Publishing: Cham, Switzerland, 2018; pp. 451–455. [Google Scholar] [CrossRef]
  29. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
  30. Zou, H.; Hastie, T. Regularization and variable selection via the elastic net. J. R. Stat. Soc. Ser. B Stat. Methodol. 2005, 67, 301–320. [Google Scholar] [CrossRef]
  31. Breiman, L.; Friedman, J.; Olshen, R.A.; Stone, C.J. Classification and Regression Trees; Chapman and Hall/CRC: Boca Raton, FL, USA, 2017. [Google Scholar] [CrossRef]
  32. Breiman, L. Bagging predictors. Mach. Learn. 1996, 24, 123–140. [Google Scholar] [CrossRef]
  33. Denisko, D.; Hoffman, M.M. Classification and interaction in random forests. Proc. Natl. Acad. Sci. USA 2018, 115, 1690–1692. [Google Scholar] [CrossRef]
  34. Yin, L.; Lin, S.; Sun, Z.; Wang, S.; Li, R.; He, Y. PriMonitor: An adaptive tuning privacy-preserving approach for multimodal emotion detection. World Wide Web 2024, 27, 9. [Google Scholar] [CrossRef]
  35. Yin, L.; Lin, S.; Sun, Z.; Li, R.; He, Y.; Hao, Z. A game-theoretic approach for federated learning: A trade-off among privacy, accuracy and energy. Digit. Commun. Netw. 2024, 10, 389–403. [Google Scholar] [CrossRef]
  36. Shi, S.; Han, D.; Cui, M. A multimodal hybrid parallel network intrusion detection model. Connect. Sci. 2023, 35, 2227780. [Google Scholar] [CrossRef]
  37. Raschka, S.; Patterson, J.; Nolet, C. Machine learning in python: Main developments and technology trends in data science, machine learning, and artificial intelligence. Information 2020, 11, 193. [Google Scholar] [CrossRef]
  38. Hancock, J.T.; Khoshgoftaar, T.M. CatBoost for big data: An interdisciplinary review. J. Big Data 2020, 7, 94. [Google Scholar] [CrossRef]
  39. Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. Catboost: Unbiased boosting with categorical features. Adv. Neural Inf. Process. Syst. 2018, 31, 6638–6648. [Google Scholar] [CrossRef]
  40. Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A highly efficient gradient boosting decision tree. Adv. Neural Inf. Process. Syst. 2017, 30, 3149–3157. [Google Scholar] [CrossRef]
  41. Ju, Y.; Sun, G.; Chen, Q.; Zhang, M.; Zhu, H.; Rehman, M.U. A model combining convolutional neural network and Light-GBM algorithm for ultra-short-term wind power forecasting. IEEE Access 2019, 7, 28309–28318. [Google Scholar] [CrossRef]
  42. Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 4–8 August 2019; pp. 2623–2631. [Google Scholar] [CrossRef]
  43. Bergstra, J.; Bardenet, R.; Bengio, Y.; Kégl, B. Algorithms for hyper-parameter optimization. Adv. Neural Inf. Process. Syst. 2011, 24. [Google Scholar] [CrossRef]
  44. Bergstra, J.; Bengio, Y. Random search for hyper-parameter optimization. J. Mach. Learn. Res. 2012, 13, 281–305. [Google Scholar] [CrossRef]
  45. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. (NeurIPS) 2017, 30, 4765–4774. [Google Scholar] [CrossRef]
  46. Mangalathu, S.; Hwang, S.H.; Jeon, J.S. Failure mode and effects analysis of RC members based on machine-learning-based Shapley Additive exPlanations (SHAP) approach. Eng. Struct. 2020, 219, 110927. [Google Scholar] [CrossRef]
Figure 1. Distribution of the input and output features.
Figure 1. Distribution of the input and output features.
Algorithms 19 00277 g001
Figure 2. Correlation of the input and output variables.
Figure 2. Correlation of the input and output variables.
Algorithms 19 00277 g002
Figure 3. Block diagram for MLP Algorithm.
Figure 3. Block diagram for MLP Algorithm.
Algorithms 19 00277 g003
Figure 4. Block diagram for Decision Tree Algorithm.
Figure 4. Block diagram for Decision Tree Algorithm.
Algorithms 19 00277 g004
Figure 5. Block diagram for RF Algorithm.
Figure 5. Block diagram for RF Algorithm.
Algorithms 19 00277 g005
Figure 6. The iterative process of XGBoost.
Figure 6. The iterative process of XGBoost.
Algorithms 19 00277 g006
Figure 7. Predicted vs. True Diagrams for (Outer) Validation Folds of Train Subset.
Figure 7. Predicted vs. True Diagrams for (Outer) Validation Folds of Train Subset.
Algorithms 19 00277 g007
Figure 8. Predicted vs. True Diagrams for Holdout Subset.
Figure 8. Predicted vs. True Diagrams for Holdout Subset.
Algorithms 19 00277 g008
Figure 9. CV Validation R2 for each Outer Validation Fold of Train Subset.
Figure 9. CV Validation R2 for each Outer Validation Fold of Train Subset.
Algorithms 19 00277 g009
Figure 10. Feature Importance Analysis Outputs for Best-Performing Models.
Figure 10. Feature Importance Analysis Outputs for Best-Performing Models.
Algorithms 19 00277 g010
Figure 11. SHAP Summary Plot for XGBoost Model.
Figure 11. SHAP Summary Plot for XGBoost Model.
Algorithms 19 00277 g011
Figure 12. Learning Curves for ML Models.
Figure 12. Learning Curves for ML Models.
Algorithms 19 00277 g012
Table 1. Hyperparameter Search Spaces.
Table 1. Hyperparameter Search Spaces.
ModelParameterTypeRangeScale
XGBoostn_estimatorsint[100, 800]linear
learning_ratefloat[0.01, 0.3]log
max_depthint[3, 10]linear
subsamplefloat[0.5, 1.0]linear
colsample_bytreefloat[0.5, 1.0]linear
CatBoostiterationsint[100, 800]linear
learning_ratefloat[0.01, 0.3]log
depthint[3, 10]linear
LightGBMn_estimatorsint[100, 800]linear
learning_ratefloat[0.01, 0.3]log
max_depthint[3, 10]linear
subsamplefloat[0.5, 1.0]linear
colsample_bytreefloat[0.5, 1.0]linear
ElasticNetalphafloat[1 × 10−4, 10.0]log
l1_ratiofloat[0.0, 1.0]linear
max_iter-5000fixed
Decision Treemax_depthint[2, 15]linear
min_samples_leafint[1, 20]linear
Random Forestn_estimatorsint[100, 600]linear
max_featurescategorical{“sqrt”, “log2”, 0.5, 0.7}-
Bagging (DT base)n_estimatorsint[50, 300]linear
base_max_depthint[3, 12]linear
KNNn_neighborsint[3, 20]linear
weightscategorical{“uniform”, “distance”}-
MLPn_layersint[1, 3]linear
units_l{i}int[32, 256]linear
activationcategorical{“relu”, “tanh”}-
learning_rate_initfloat[1 × 10−4, 1 × 10−2]log
solver-adamfixed
max_iter-1000fixed
early_stopping-Truefixed
validation_fraction-0.1fixed
Table 2. Nested CV Structure.
Table 2. Nested CV Structure.
LevelTypeFoldsShufflePurpose
Holdout splittrain_test_splitYes (seed 42)20% final holdout set, never seen during tuning
Outer loopKFold10Yes (seed 42)Unbiased performance estimation
Inner loopKFold5Yes (seed 42)Optuna hyperparameter optimization objective
Table 3. Explanation of Nested CV Metrics.
Table 3. Explanation of Nested CV Metrics.
ColumnExplanation
CV Train R2Mean R2 on the outer-train fold across 10 folds of data the model was just trained on
CV Train RMSE/MAEMean train errors across 10 outer folds
CV Val R2Mean R2 on the unseen outer-val fold across 10 folds—the realistic CV estimate
CV Val RMSE/MAEMean validation errors on the unseen outer-val fold across 10 folds—the realistic CV estimate
Hold R2Holdout R2–final performance
Hold RMSE/MAEHoldout errors–final performance
Table 4. Explanation of Permutation Significance Test Metrics.
Table 4. Explanation of Permutation Significance Test Metrics.
MetricDefinition
Real Hold R2Coefficient of determination (R2) obtained from the model trained on the original (unshuffled) dataset and evaluated on the independent holdout set; represents the observed model performance T real .
Permuted R2 MeanMean R2 across all permutation runs, where each value corresponds to T i perm , i.e., representing the R2 obtained after randomly shuffling the target variable, retraining the model from scratch under the same hyperparameter settings, and evaluating it on the unchanged holdout set; this metric represents the expected model performance under the null hypothesis of no true relationship between the input features and the target variable.
Permuted R2 StdStandard deviation of R2 across permutation runs; this quantifies the variability of the null distribution.
Permuted R2 MaxMaximum R2 observed among all permutation runs; this represents the most extreme performance achievable by chance under the null hypothesis.
p-valueEmpirical probability that a model trained on randomly permuted targets achieves performance equal to or greater than the real model performance, computed as p = (1 + k)/(N + 1), where k is the number of permutation runs satisfying R2_perm ≥ R2_real and N is the total number of permutations.
SignificanceIndicator of statistical significance based on the permutation test; the result is considered statistically significant if p < 0.05, indicating that the observed performance is unlikely to arise from random chance.
Table 5. Optimized Hyperparameters.
Table 5. Optimized Hyperparameters.
ModelOptimized Hyperparameter Values
XGBoost{‘n_estimators’: 507, ‘learning_rate’: 0.0835, ‘max_depth’: 4, ‘subsample’: 0.6675, ‘colsample_bytree’: 0.7838}
CatBoost{‘iterations’: 703, ‘learning_rate’: 0.1686, ‘depth’: 5}
LightGBM{‘n_estimators’: 642, ‘learning_rate’: 0.0873, ‘max_depth’: 9, ‘subsample’: 0.9098, ‘colsample_bytree’: 0.6368}
ElasticNet{‘alpha’: 0.0203, ‘l1_ratio’: 0.9966}
DecisionTree{‘max_depth’: 14, ‘min_samples_leaf’: 4}
RandomForest{‘n_estimators’: 236, ‘max_features’: 0.7}
Bagging{‘n_estimators’: 241, ‘base_max_depth’: 12}
KNN{‘n_neighbors’: 4, ‘weights’: ‘distance’}
MLP{‘hidden_layer_sizes’: (214, 241, 245), ‘activation’: ‘relu’, ‘learning_rate_init’: 0.0003513}
Table 6. Evaluation Metrics of ML Models (Nested CV Outer Folds and Holdout Set).
Table 6. Evaluation Metrics of ML Models (Nested CV Outer Folds and Holdout Set).
ModelCV Train R2CV Train RMSECV Val R2CV Val RMSEHoldout R2Holdout RMSE
XGBoost0.98961.77140.86666.41510.83987.0867
CatBoost0.99041.71360.87926.13670.83677.1555
LightGBM0.99001.76290.86516.46020.82397.4300
Elastic Net0.592211.43510.573811.51210.597711.2312
Decision Tree0.90895.16790.71469.42990.71569.4424
Random Forest0.97572.79120.8337.23940.81217.6756
Bagging0.96533.33770.81257.65250.79757.9679
KNN0.99790.82290.78948.15780.75238.8132
MLP0.89985.64060.80197.85350.77918.3224
Table 7. Permutation Significance Test Results.
Table 7. Permutation Significance Test Results.
ModelReal Hold R2Perm R2 MeanPerm R2 StdPerm R2 Maxp-ValueSignificant
XGBoost0.8398−0.34030.1095−0.1676<<0.0001Yes
CatBoost0.8367−0.3070.1075−0.0503<<0.0001Yes
LightGBM0.8239−0.46960.1137−0.1622<<0.0001Yes
Elastic Net0.5977−0.00890.04430.1017<<0.0001Yes
Decision Tree0.7156−0.46160.1242−0.1618<<0.0001Yes
Random Forest0.8121−0.20710.0767−0.0286<<0.0001Yes
Bagging0.7975−0.1160.07130.0633<<0.0001Yes
KNN0.7523−0.41620.0932−0.1786<<0.0001Yes
MLP0.7791−0.10230.06860.08<<0.0001Yes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Işıkdağ, Ü.; Bekdaş, G.; Nigdeli, S.M.; Ahadian, F.; Geem, Z.W. Robust Prediction of Compressive Strength of SCM Concrete with Nested Cross-Validation and Bayesian Optimization. Algorithms 2026, 19, 277. https://doi.org/10.3390/a19040277

AMA Style

Işıkdağ Ü, Bekdaş G, Nigdeli SM, Ahadian F, Geem ZW. Robust Prediction of Compressive Strength of SCM Concrete with Nested Cross-Validation and Bayesian Optimization. Algorithms. 2026; 19(4):277. https://doi.org/10.3390/a19040277

Chicago/Turabian Style

Işıkdağ, Ümit, Gebrail Bekdaş, Sinan Melih Nigdeli, Farnaz Ahadian, and Zong Woo Geem. 2026. "Robust Prediction of Compressive Strength of SCM Concrete with Nested Cross-Validation and Bayesian Optimization" Algorithms 19, no. 4: 277. https://doi.org/10.3390/a19040277

APA Style

Işıkdağ, Ü., Bekdaş, G., Nigdeli, S. M., Ahadian, F., & Geem, Z. W. (2026). Robust Prediction of Compressive Strength of SCM Concrete with Nested Cross-Validation and Bayesian Optimization. Algorithms, 19(4), 277. https://doi.org/10.3390/a19040277

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop