Next Article in Journal
Statement of Peer Review
Previous Article in Journal
A Comparative Assessment of XFEM and FEM for Stress Concentration at Circular Holes near Bi-Material Interfaces
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Proceeding Paper

Key Predictors of Lightweight Aggregate Concrete Compressive Strength by Machine Learning from Density Parameters and Ultrasonic Pulse Velocity Testing †

by
Violeta Migallón
,
Héctor Penadés
* and
José Penadés
Department of Computer Science and Artificial Intelligence, University of Alicante, 03690 Alicante, Spain
*
Author to whom correspondence should be addressed.
Presented at the 4th International Online Conference on Materials, 3–6 November 2025; Available online: https://sciforum.net/event/IOCM2025.
Mater. Proc. 2025, 26(1), 4; https://doi.org/10.3390/materproc2025026004
Published: 6 January 2026
(This article belongs to the Proceedings of The 4th International Online Conference on Materials)

Abstract

Non-destructive evaluation techniques are increasingly recognised as effective alternatives to destructive testing for estimating the compressive strength of lightweight aggregate concrete (LWAC). Among these, ultrasonic pulse velocity (UPV) is a well-established and widely employed method, characterised by its speed, non-invasiveness, and relative simplicity of implementation. In this study, an experimental dataset comprising 640 core segments from 160 cylindrical specimens, provided for analysis, was investigated. Each segment was described by physical and processing variables or features, including lightweight aggregate (LWA) and concrete densities, casting and vibration times, experimental dry density, and P-wave velocity obtained through UPV testing. A segregation index, derived from UPV measurements and defined as the ratio of local to mean P-wave velocity within each specimen, was also considered, following approaches previously suggested in the literature. A range of machine learning techniques was applied to assess the predictive capacity of local P-wave velocity and segregation index. Most ensemble-based methods and support vector regression (SVR) achieved the highest predictive performance when the segregation index was excluded, suggesting that its inclusion did not improve the predictive ability of the models. By contrast, Gaussian process regression (GPR) showed slight improvements when the segregation index was included. The results confirmed that the P-wave velocity measured by UPV testing is a reliable non-destructive predictor of compressive strength in LWAC. At the same time, the added value of the segregation index remained negligible under conditions of low segregation, as reflected by segregation index values above 0.8. These findings highlight the practical potential of integrating UPV-based measurements with data-driven modelling to enhance the reliability of concrete characterisation and quality control.

1. Introduction

Recent trends in sustainable and efficient construction have increased interest in lightweight aggregate concrete (LWAC). This material allows engineers to design lighter structures with improved thermal performance and reduced seismic demand, making it a valuable alternative to traditional concrete in both structural and non-structural applications. However, the benefits of LWAC depend strongly on the uniform distribution of its components. During casting and vibration, differences in particle density can cause the lightweight aggregates (LWAs) to rise while heavier materials settle, producing segregation. As a result, variations in density and porosity appear along the element, reducing strength and compromising long-term durability [1]. Preventing this behaviour requires controlling the fresh-state properties of the mixture to achieve sufficient workability without segregation. Proper rheology and carefully applied compaction are essential to maintaining a homogeneous structure [2].
Segregation in LWAC can significantly influence its mechanical performance and durability, and therefore must be carefully evaluated. Several experimental methods have been proposed to quantify this phenomenon [3,4,5,6]. Density-variation tests assess changes in volumetric mass along the height of a specimen, while image-analysis procedures enable the spatial distribution of aggregates to be observed in polished cross-sections. Among non-destructive techniques, the ultrasonic pulse velocity (UPV) method is particularly suitable for assessing internal homogeneity and detecting voids or microcracks. UPV measurements are based on the travel time of an acoustic wave through the concrete; lower velocities indicate discontinuities or higher porosity, while higher velocities correspond to denser and more uniform regions. Due to its simplicity and non-invasive nature, UPV is widely used for quality control in both laboratory and field applications, providing an effective means of identifying segregation or internal defects before they compromise structural performance. Nevertheless, establishing reliable correlations between UPV results, density, and compressive strength in LWAC is not straightforward. The heterogeneous and highly porous nature of LWAs introduces complex interactions among vibration time, water absorption, density, and wave propagation. These interdependencies limit the accuracy of traditional empirical or regression-based models.
Recent studies have adopted data-driven and machine learning approaches to model the behaviour of concrete and to predict its mechanical properties more accurately [7,8,9]. Methods such as artificial neural networks (ANNs), support vector regression (SVR), and tree-based ensemble models have shown promising results in capturing multivariable interactions that traditional empirical models cannot effectively represent [10,11,12,13,14,15,16,17]. ANNs, particularly multilayer perceptrons (MLPs), have been extensively used to model the compressive strength of various types of concrete. Several studies have demonstrated the superior accuracy of these networks compared to traditional statistical models, such as multiple linear regression (MLR) [10,11]. Similarly, it has been reported that ANNs outperform both non-linear regression and regression tree (RT) models in predicting the 28-day compressive strength of recycled aggregate concrete [13]. ANN models have also been applied to LWAC. In [14], an MLP trained with the Levenberg–Marquardt algorithm was used to predict LWAC compressive strength from variables such as particle density, vibration time, and P-wave velocity. The same dataset was later analysed in [16], where ANN and regression tree models outperformed MLR. More recent analyses [17] showed that ensemble-based methods outperformed ANNs and achieved the best predictive performance. Gaussian process regression (GPR) has also shown strong potential. In [18], GPR outperformed SVR, ANN, RT, and ensemble models on a small LWAC dataset. Likewise, [19] evaluated several machine learning algorithms on a larger compiled dataset and found that GPR was among the most accurate models for predicting LWAC compressive strength.
In this context, the present study extends the previous work [17] by restricting data normalisation to the training set, incorporating additional machine learning models, and performing a more systematic analysis of the relevance of input variables through variable selection, thereby refining the predictive framework and improving the interpretability of the model. It explores the use of machine learning techniques combined with UPV data to predict the compressive strength of LWAC. The influence of the P-wave velocity and the segregation index on model performance is assessed, and the performance of different algorithms is compared to identify the most effective approach for capturing the complex, non-linear behaviour of LWAC. The outcomes aim to contribute to the development of reliable, non-destructive tools for quality control and mix-design optimisation in lightweight concrete systems.

2. Materials and Methods

2.1. Experimental Dataset

The dataset used in this study, described in detail in [14], was obtained by the Materials and Territory Technology Research Group at the University of Alicante (Spain). It comprises 640 core segments extracted from 160 cylindrical specimens of LWAC. The specimens were produced using four mix designs with different densities and aggregate types, including expanded clay as the lightweight coarse aggregate and a natural fine limestone aggregate as the fine fraction. The proportions of the mixture were determined following the Fanjul method [20], a procedure specifically developed for LWAC mix design. Aggregate properties such as bulk density, particle density, and water absorption were characterised according to Spanish and European standards (UNE EN 1097-3 [21], UNE EN 1097-6 [22] and UNE EN 933-1 [23]). The four LWAC types resulted from combining two target concrete densities (1700 and 1900 kg/m3) with two expanded-clay LWAs of different particle densities (482 and 1019 kg/m3), corresponding to granulometric fractions 6/10 and 4/10, respectively. From each cylindrical specimen, two 50 mm diameter cores were extracted; however, only one core from each specimen was selected for testing. This selected core was then cut into four 70 mm long segments, resulting in the 640 segments used in the dataset.
To analyse the effect of segregation, vibration time during compaction was intentionally varied between 0 and 80 s, while four concrete laying times (15, 30, 60, and 90 min) were applied to reproduce realistic construction conditions. To evaluate internal non-uniformity, a segregation index was defined based on UPV measurements taken along each concrete core. After 28 days of curing, each core was divided into four equal segments, and the P-wave velocity ( V P i ) was measured for each one. Velocity measurements were obtained using through-transmission, with compressional waves recorded by 250 kHz Panametrics transducers. The segregation index was calculated as the ratio between the P-wave velocity measured in a single segment i and the mean P-wave velocity of all segments from the same specimen ( V P ¯ ), thereby enabling its calculation for each segment, as follows:
S I i = V P i V P ¯ , i = 1 , 2 , 3 , 4 .
Values of this index close to one indicate a homogeneous material with no significant differences in density. Values greater than one correspond to denser and stiffer regions, whereas values lower than one represent lighter and more porous zones within the specimen. These variations reflect local differences in compactness caused by the upward or downward movement of LWAs during casting and vibration.
The target variable in this study was the compressive strength of LWAC. The predictor variables included parameters describing the mixture composition, production process and physical properties. Two density-related variables were considered: the fixed density of LWAC (kg/m3), which defines the intended mixture compactness, and the particle density of LWA (kg/m3), which reflects the intrinsic characteristics of the aggregate. In addition, the dataset contained the concrete laying time (min), vibration time (s), experimental dry density (kg/m3), P-wave velocity (m/s), and the segregation index obtained as described previously. The compressive strength values in the dataset ranged from 2.99 to 50.72 MPa, with a mean of 21.55 MPa and a median of 20.25 MPa (see Figure 1). A summary of descriptive statistics for all input and target variables is provided in Table 1.
Figure 2 shows the Spearman correlation heatmap for the LWAC dataset used to investigate potential interaction effects among the predictors, as the normality assumption was not satisfied for any of the variables. Compressive strength exhibits strong positive monotonic associations with the particle density of LWA ( ρ = 0.71 ) and experimental dry density ( ρ = 0.65 ). LWAC fixed density also correlates positively with experimental dry density ( ρ = 0.65 ) and P-wave velocity ( ρ = 0.52 ). A moderate inverse relationship is observed between P-wave velocity and LWA particle density ( ρ = 0.40 ). In contrast, the time-related variables (concrete laying time and vibration time) show negligible correlations with the remaining variables.

2.2. Machine Learning Workflow

To model the relationships between experimental variables and the compressive strength of LWAC, several machine learning algorithms were implemented and compared. The selected models represent a broad range of regression paradigms, ranging from simple non-parametric approaches such as K-nearest neighbours (KNNs), to neural and kernel-based models, including ANNs and SVR, as well as tree-based ensemble and probabilistic methods such as random forest (RF), gradient boosting regressor (GBR), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), and GPR, as well as hybrid ensemble techniques (HETs) combining heterogeneous base learners. The selection of algorithms was based on their widespread use in concrete property prediction and on the fact that different modelling families respond in distinct ways to non-linearity and data heterogeneity. Non-parametric methods can capture local patterns directly from the data, whereas kernel-based approaches such as SVR are well suited to modelling smooth non-linear relationships. Tree-based ensemble methods are particularly robust in the presence of variable interactions and noise, and have repeatedly shown strong performance in prediction tasks [7]. Probabilistic techniques such as GPR provide uncertainty estimates that are useful for assessing the reliability of predictions. Hybrid ensembles combine complementary learners to enhance accuracy, although their performance can overlap with that of strong individual models and they may require substantially higher computational costs. Each modelling family also presents inherent limitations. For example, non-parametric and kernel-based methods can be sensitive to variable scaling; ANNs and GPR depend strongly on hyperparameter tuning; and tree-based ensembles may be prone to overfitting. These aspects may influence model performance depending on the structure and heterogeneity of the dataset. Including this range of methods provides a broad and representative comparison of modelling strategies for predicting LWAC compressive strength.
All machine learning models were developed following a systematic workflow that included data preprocessing, model training, optimisation, and performance evaluation. Prior to model training, several data preprocessing and validation steps were carried out to ensure reliable and reproducible results. Since the experimental dataset contained no missing values, data preparation was straightforward. The dataset was randomly divided into two subsets, with 75% of the samples used for training and 25% for testing. All input variables were scaled using the Min-Max normalisation method, which rescales each variable to the [0,1] range based on its minimum and maximum values from the training subset. The same transformation was then applied to the test data to avoid information leakage. No additional preprocessing steps were applied beyond this feature scaling. This partitioning procedure was repeated 10 times with different random splits, following a Monte Carlo cross-validation scheme, to ensure that the results were consistent and not dependent on a single division of the data. To ensure reproducibility, fixed random seeds were used, allowing the same partitions to be generated in future analyses.
A preliminary analysis was performed to identify suitable hyperparameter ranges for each algorithm, followed by a grid search procedure to fine-tune the models and determine the best configuration for each case. Grid search was carried out exclusively on the training subset, using internal cross-validation to prevent information leakage during the optimisation process. All algorithms were trained and evaluated under the same conditions and using identical input variables to ensure a consistent comparison.
Model performance was assessed using three widely applied statistical indicators: the coefficient of determination ( R 2 ), the root mean square error (RMSE), and the mean absolute error (MAE). The R 2 value measures how well the model explains the variability of the observed data, with values closer to one indicating a stronger predictive performance. RMSE and MAE represent the average magnitude of prediction errors, although RMSE gives more weight to larger deviations.
All models were implemented in Python 3.12.7, using Scikit-learn 1.6.1 [24] as the main library for machine learning, and the XGBoost [25] and LightGBM [26] libraries for the gradient boosting models. The experiments were executed in a Windows 10 (64-bit) environment, running on a laptop equipped with an Intel Core i7-1065G7 processor (1.30 GHz), 16 GB of RAM, and an NVIDIA GeForce RTX 3050 GPU (4 GB GDDR6 VRAM).

3. Numerical Results and Discussion

Table 2 presents the average predictive performance of all machine learning models on the test datasets, computed over 10 independent random partitions. The overall results reveal differences in predictive capability across the various modelling approaches. Among the tested algorithms, the KNN, ANN, and GPR models exhibited comparatively lower performance with R 2 values in the range of 0.7868–0.8172 and higher RMSE values. For these models, the exclusion of the segregation index led to a slight deterioration in predictive performance, suggesting that this variable provides useful complementary information to that obtained from the P-wave velocity, whereas no such effect was observed for other models.
The SVR, tree-based, and hybrid ensemble methods (RF, GBR, XGBoost, LightGBM, and HET) achieved higher performance, with GBR and HET performing best among them. When the segregation index was included, GBR reached an R 2 of 0.8279, an RMSE of 3.7079 MPa, and an MAE of 2.8962 MPa. Excluding this variable led to a slight improvement, with an R 2 of 0.8291, an RMSE of 3.6943 MPa, and an MAE of 2.8775 MPa. Among the hybrid ensembles tested, the weighted combination of SVR, XGBoost, and RF achieved a higher overall performance than the individual models, whereas other configurations did not provide further improvement. With the segregation index included, it obtained an R 2 of 0.8271, an RMSE of 3.7109 MPa, and an MAE of 2.8777 MPa. When this index was excluded, the HET model showed a slight improvement, achieving an R 2 of 0.8320, an RMSE of 3.6575 MPa, and an MAE of 2.8361 MPa. Figure 3 shows the relationship between the observed and predicted compressive strength values for the GBR and HET models without the segregation index. The slight improvement observed after excluding the segregation index can be attributed to its limited variability in the experimental dataset (0.845–1.136), which reduces its contribution to model performance compared with other, more informative predictors such as P-wave velocity.
To evaluate the statistical robustness of the predictive performance, an additional analysis was conducted focusing on the two best-performing models (GBR and HET). Table 3 reports the 95% confidence intervals (CIs) for R 2 , RMSE, and MAE, obtained through non-parametric bootstrap resampling (5000 iterations applied to the 10 Monte Carlo test partitions). The overlapping intervals observed for all three metrics suggest that the differences between the two models are small relative to their sampling variability. This interpretation is supported by the Wilcoxon signed-rank test, which yielded p-values above 0.05 for all metrics. Consequently, the performance differences between GBR and HET cannot be considered statistically significant.
To assess the contribution of each input variable to the model predictions, a permutation-based feature importance analysis was performed. This approach quantifies the decrease in model performance when the values of a given variable are randomly permuted, thereby breaking its relationship with the target variable. The greater the reduction in performance, the more influential that feature is considered to be. Figure 4 shows the feature importance rankings obtained for the GBR and HET models. The overall pattern is similar for both algorithms, with LWA particle density and experimental dry density clearly emerging as the dominant predictors. These results confirm that density-related parameters are the key factors controlling the compressive strength of LWAC. The ranking of the less influential variables differs slightly between the two models. In the GBR model, P-wave velocity showed a somewhat higher contribution (5.5%) compared with the HET model (2.6%), indicating that UPV-based measurements still provide useful complementary information about the internal structure of the material. Although less dominant than the density-related parameters, P-wave velocity remains a reliable non-destructive indicator that helps refine the prediction of compressive strength.
On the other hand, interaction effects were first quantified using the H-statistic of Friedman and Popescu [27], which provides a global measure of the extent to which pairs of predictors jointly contribute to the model response. Complementarily, conditional 3D partial dependence plots [28] were used to visualise how pairs of predictors influence compressive strength. For each predictor pair, values were varied over a predefined grid while all remaining variables were held at their observed values. At every grid point, the model was evaluated across the conditioned dataset, and the resulting surface value at that point represents the average of these conditioned predictions. Both the GBR and HET models exhibited essentially the same interaction structure, with differences limited to minor local variations, and with HET producing smoother response surfaces due to its ensemble nature. The response was driven predominantly by interactions, with H-statistic values exceeding 0.87 across variables. The four strongest interaction effects (each with H 1 ) involved the following predictor pairs: LWA particle density with vibration time, LWA particle density with concrete laying time, LWA fixed density with LWA particle density, and LWA particle density with P-wave velocity. In the case of LWA particle density and vibration time, increases in both predictors corresponded to higher predicted compressive strength, with the effect of vibration time becoming more evident at higher particle densities (see Figure 5a). A similarly clear joint effect was observed for LWA particle density and P-wave velocity, with the highest responses occurring when both variables were in their upper ranges (see Figure 5b). The remaining two interaction pairs showed no qualitatively different patterns.
The computational efficiency of each model was evaluated by measuring the execution time per random seed, using a single fixed hyperparameter configuration. The results, summarised in Table 4, show that the SVR model was the fastest, with a mean execution time of only 0.016 s, while GPR required the longest time (1.59 s). All results correspond to averages over 10 random seeds, with parallelisation enabled whenever supported by the underlying library. Certain algorithms (e.g., SVR and GPR) were executed in single-thread mode due to the lack of native parallel support.
Figure 6 illustrates how the mean execution time increases with the number of hyperparameter combinations explored during grid search. Each grid configuration represents a distinct parameter set, allowing for the assessment of how the computational cost scales with model complexity. Among the best-performing algorithms, the GBR and the hybrid ensemble model achieved comparable predictive performance but differed substantially in computational cost. For instance, a grid search with 24 hyperparameter combinations required approximately 14 s for GBR, whereas the HET model took nearly 3 min to complete the same process across 10 runs with different random seeds. This clear difference highlights the higher computational burden of the hybrid approach, resulting from the integration of multiple base learners. Therefore, GBR offers a more favourable balance between performance and computational efficiency, making it the most suitable option when both speed and precision are required.

4. Conclusions

The present study evaluated the effectiveness of machine learning techniques to predict the compressive strength of LWAC using UPV and mix-related variables. Among the tested algorithms, GBR and a hybrid ensemble model achieved the highest predictive performance, with mean R 2 values around 0.83. Although GBR exhibited the best individual performance, the weighted hybrid ensemble combining SVR, XGBoost, and RF achieved slightly higher predictive accuracy. This may be due to the complementary nature of the base models, with XGBoost introducing additional regularisation and subsampling that enhanced diversity within the ensemble. On the other hand, the HET model required substantially more computational time than the GBR model. The analysis of feature importance revealed that density-related parameters, particularly the LWA particle density and the experimental dry density, were the dominant predictors of compressive strength. The P-wave velocity obtained from UPV testing contributed additional, but secondary, information about the internal structure of the material. The segregation index did not meaningfully contribute to the predictive models under the present conditions of low segregation and specific LWAC mixes, and in most cases slightly reduced their predictive accuracy.
From a practical standpoint, the modelling approach developed in this study may provide a useful supplementary tool for estimating compressive strength from non-destructive measurements and basic mix parameters. These estimates could assist preliminary evaluations within quality-control processes or complement standard laboratory testing. Although the present work was limited to the variables and material types available in the original dataset, additional process-related and material-related factors, such as curing conditions, moisture content, mixture proportions or aggregate grading, may also influence the compressive strength of LWAC and could be incorporated in future research to develop more comprehensive and robust predictive models.
Future work will explore the application of deep learning combined with explainable AI to improve predictive accuracy and interpretability. Furthermore, expanding the dataset to include more diverse mixtures, lightweight aggregates and production environments would help strengthen model generalisation and reliability, thereby helping to improve industry applicability.

Author Contributions

Conceptualisation, V.M., H.P. and J.P.; methodology, V.M., H.P. and J.P.; software, V.M., H.P. and J.P.; validation, V.M., H.P. and J.P.; formal analysis, V.M., H.P. and J.P.; investigation, V.M., H.P. and J.P.; data curation, V.M., H.P. and J.P.; writing—original draft preparation, V.M., H.P. and J.P.; writing—review and editing, V.M., H.P. and J.P.; supervision, V.M. and J.P. All authors have read and agreed to the published version of the manuscript.

Funding

This research was partially funded by MCIN/AEI/10.13039/501100011033, grant PID2021-123627OB-C55, and by ‘ERDF A way of making Europe’.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were generated.

Acknowledgments

We thank Antonio J. Tenza-Abril, leader of the Materials and Territory Technology Research Group at the University of Alicante, for providing the data used in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Chandra, S.; Berntsson, L. Lightweight Aggregate Concrete; Elsevier: Amsterdam, The Netherlands, 2002. [Google Scholar]
  2. Tenza-Abril, A.; Benavente, D.; Pla, C.; Baeza-Brotons, F.; Valdes-Abellan, J.; Solak, A. Statistical and experimental study for determining the influence of the segregation phenomenon on physical and mechanical properties of lightweight concrete. Constr. Build. Mater. 2020, 238, 117642. [Google Scholar] [CrossRef] [Scilit]
  3. Navarrete, I.; Lopez, M. Estimating the segregation of concrete based on mixture design and vibratory energy. Constr. Build. Mater. 2016, 122, 384–390. [Google Scholar] [CrossRef] [Scilit]
  4. Solak, A.M.; Tenza-Abril, A.J.; Baeza-Brotons, F.; Benavente, D. Proposing a New Method Based on Image Analysis to Estimate the Segregation Index of Lightweight Aggregate Concretes. Mater. 2019, 12, 3642. [Google Scholar] [CrossRef] [Scilit]
  5. Solak, A.M.; Tenza-Abril, A.J.; García-Vera, V.E. Adopting an image analysis method to study the influence of segregation on the compressive strength of lightweight aggregate concretes. Constr. Build. Mater. 2022, 323, 126594. [Google Scholar] [CrossRef] [Scilit]
  6. Ahmed, H.; Punkki, J. Methods for assessing concrete segregation due to compaction. Nord. Concr. Res 2024, 70, 1–23. [Google Scholar] [CrossRef] [Scilit]
  7. Ni, B.; Rahman, M.Z.; Guo, S.; Zhu, D. A review on properties and multi-objective performance predictions of concrete based on machine learning models. Mater. Today Commun. 2025, 44, 112017. [Google Scholar] [CrossRef] [Scilit]
  8. Behera, D.; Liu, K.Y.; Rachman, F.; Worku, A.M. Innovations and Applications in Lightweight Concrete: Review of Current Practices and Future Directions. Buildings 2025, 15, 2113. [Google Scholar] [CrossRef] [Scilit]
  9. Adsul, N.; Choi, Y.; Kang, S.T. A Comprehensive Review of Numerical and Machine Learning Approaches for Predicting Concrete Properties: From Fresh to Long-Term. Materials 2025, 18, 3718. [Google Scholar] [CrossRef] [Scilit]
  10. Kewalramani, M.A.; Gupta, R. Concrete compressive strength prediction using ultrasonic pulse velocity through artificial neural networks. Automat. Constr. 2006, 15, 374–379. [Google Scholar] [CrossRef] [Scilit]
  11. Tavakkol, S.; Alapour, F.; Kazemian, A.; Hasaninejad, A.; Ghanbari, A.; Ramezanianpour, A.A. Prediction of lightweight concrete strength by categorized regression, MLR and ANN. Comput. Concr. 2013, 12, 151–167. [Google Scholar] [CrossRef] [Scilit]
  12. Charhate, S.; Subhedar, M.; Adsul, N. Prediction of Concrete Properties Using Multiple Linear Regression and Artificial Neural Network. J. Soft Comput. Civ. Eng. 2018, 2, 27–38. [Google Scholar] [CrossRef] [Scilit]
  13. Deshpande, N.; Londhe, S.; Kulkarni, S. Modeling compressive strength of recycled aggregate concrete by Artificial Neural Network, Model Tree and Non-linear Regression. Int. J. Sustain. Built Environ. 2014, 3, 187–198. [Google Scholar] [CrossRef] [Scilit]
  14. Tenza-Abril, A.J.; Villacampa, Y.; Solak, A.M.; Baeza-Brotons, F. Prediction and sensitivity analysis of compressive strength in segregated lightweight concrete based on artificial neural network using ultrasonic pulse velocity. Constr. Build. Mater. 2018, 189, 1173–1183. [Google Scholar] [CrossRef] [Scilit]
  15. Sajan, K.C.; Bhusal, A.; Gautam, D.; Rupakhety, R. Earthquake damage and rehabilitation intervention prediction using machine learning. Eng. Fail. Anal. 2023, 144, 106949. [Google Scholar] [CrossRef] [Scilit]
  16. Migallón, V.; Navarro-González, F.J.; Penadés, J.; Villacampa, Y. Parallel approach of a Galerkin-based methodology for predicting the compressive strength of the lightweight aggregate concrete. Constr. Build. Mater. 2019, 219, 56–68. [Google Scholar] [CrossRef] [Scilit]
  17. Migallón, V.; Penadés, H.; Penadés, J.; Tenza-Abril, A.J. A machine learning approach to prediction of the compressive strength of segregated lightweight aggregate concretes using ultrasonic pulse velocity. Appl. Sci. 2023, 13, 1953. [Google Scholar] [CrossRef] [Scilit]
  18. Kumar, A.; Arora, H.C.; Kapoor, N.R.; Mohammed, M.A.; Kumar, K.; Majumdar, A.; Thinnukool, O. Compressive Strength Prediction of Lightweight Concrete: Machine Learning Models. Sustainability 2022, 14, 2404. [Google Scholar] [CrossRef] [Scilit]
  19. Hussain, F.; Ali Khan, S.; Khushnood, R.A.; Hamza, A.; Rehman, F. Machine Learning-Based Predictive Modeling of Sustainable Lightweight Aggregate Concrete. Sustainability 2023, 15, 641. [Google Scholar] [CrossRef] [Scilit]
  20. Fernández-Fanjul, A.; Tenza-Abril, A.J. Méthode Fanjul: Dosage pondéral des bétons légers et lourds. Ann. Bâtim. Trav. Publics 2012, 5, 32–50. [Google Scholar]
  21. UNE-EN 1097-3; Tests for Mechanical and Physical Properties of Aggregates—Part 3: Determination of Loose Bulk Density and Voids. Spanish Association for Standardization: Madrid, Spain, 1999.
  22. UNE-EN 1097-6; Tests for Mechanical and Physical Properties of Aggregates—Part 6: Determination of Particle Density and Water Absorption. Spanish Association for Standardization: Madrid, Spain, 2014.
  23. UNE-EN 933-1; Tests for Geometrical Properties of Aggregates—Part 1: Determination of Particle Size Distribution—Sieving Method. Spanish Association for Standardization: Madrid, Spain, 2012.
  24. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
  25. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
  26. Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; pp. 3149–3157. [Google Scholar]
  27. Friedman, J.H.; Popescu, B.E. Predictive learning via rule ensembles. Ann. Appl. Stat. 2008, 2, 916–954. [Google Scholar] [CrossRef] [Scilit]
  28. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Scatter plot showing the compressive strength values for all lightweight aggregate concrete (LWAC) core segments.
Figure 1. Scatter plot showing the compressive strength values for all lightweight aggregate concrete (LWAC) core segments.
Materproc 26 00004 g001
Figure 2. Spearman correlation matrix. LWAC-FD: LWAC fixed density; LWA-PD: LWA particle density; CLT: Concrete laying time; VT: Vibration time; DD: Experimental dry density; PWV: P-wave velocity; SI: Segregation index; CS: Compressive strength.
Figure 2. Spearman correlation matrix. LWAC-FD: LWAC fixed density; LWA-PD: LWA particle density; CLT: Concrete laying time; VT: Vibration time; DD: Experimental dry density; PWV: P-wave velocity; SI: Segregation index; CS: Compressive strength.
Materproc 26 00004 g002
Figure 3. Observed vs. predicted compressive strength without the segregation index variable. (a) GBR model. (b) HET model.
Figure 3. Observed vs. predicted compressive strength without the segregation index variable. (a) GBR model. (b) HET model.
Materproc 26 00004 g003
Figure 4. Relative permutation importance. (a) GBR model. (b) HET model.
Figure 4. Relative permutation importance. (a) GBR model. (b) HET model.
Materproc 26 00004 g004
Figure 5. Interaction effect surfaces obtained from the conditional partial dependence plots (PDPs) for the HET model. (a) LWA particle density versus vibration time. (b) LWA particle density versus P-wave velocity.
Figure 5. Interaction effect surfaces obtained from the conditional partial dependence plots (PDPs) for the HET model. (a) LWA particle density versus vibration time. (b) LWA particle density versus P-wave velocity.
Materproc 26 00004 g005
Figure 6. Mean execution time as a function of the number of hyperparameter combinations evaluated during grid search.
Figure 6. Mean execution time as a function of the number of hyperparameter combinations evaluated during grid search.
Materproc 26 00004 g006
Table 1. Descriptive statistics of the experimental dataset.
Table 1. Descriptive statistics of the experimental dataset.
VariableMinimumMaximumMean (SD)Median (IQR)
LWAC fixed density (kg/m3)170019001800.00 (100.08)1800 (200)
LWA particle density (kg/m3)4821019750.50 (268.71)750.50 (537)
Concrete laying time (min)159048.75 (28.83)45 (64)
Vibration time (s)08030 (28.31)20 (30)
Experimental dry density (kg/m3)1069.802486.841673.35 (179.14)1677.15 (277.49)
P-wave velocity (m/s)3044.255253.733778.89 (370.88)3718.49 (425.18)
Segregation index0.8451.1361.000 (0.0352)0.999 (0.04)
Compressive strength (MPa)2.9950.7221.55 (8.97)20.25 (14.39)
SD: Standard deviation; IQR: Interquartile range; LWA: Lightweight aggregate.
Table 2. Average performance of machine learning models with and without the segregation index variable.
Table 2. Average performance of machine learning models with and without the segregation index variable.
Model R 2 R 2 (w/o SI)RMSERMSE (w/o SI)MAEMAE (w/o SI)
KNN0.78680.78444.12504.15193.17513.1885
ANN0.81710.81433.81173.83982.89172.9177
SVR0.82000.82363.78903.76692.95222.9241
RF0.82100.82243.78063.76522.93812.9154
GBR0.82790.82913.70793.69432.89622.8775
XGBoost0.82740.82673.71313.72062.89272.8979
LightGBM0.82050.82383.78563.75022.97272.9317
GPR0.81720.81603.82083.83292.96982.9768
HET0.82710.83203.71093.65752.87772.8361
All values correspond to the average performance on the test datasets across 10 independent random partitions; SI: segregation index; w/o SI: without SI. Metrics: coefficient of determination ( R 2 ), root mean square error (RMSE), and mean absolute error (MAE). Model abbreviations: KNN: K-nearest neighbour; ANN: artificial neural network; SVR: support vector regression; RF: random forest; GBR: gradient boosting regressor; XGBoost: extreme gradient boosting; LightGBM: Light gradient boosting machine; GPR: Gaussian process regression; HET: hybrid ensemble technique.
Table 3. The 95% confidence intervals (CIs) of the predictive performance metrics for the GBR and HET models (10 Monte Carlo iterations).
Table 3. The 95% confidence intervals (CIs) of the predictive performance metrics for the GBR and HET models (10 Monte Carlo iterations).
Model R 2 (95% CI)RMSE (95% CI)MAE (95% CI)
GBR0.8291 [0.8183, 0.8396]3.6943 [3.5623, 3.8262]2.8775 [2.7844, 2.9751]
HET0.8320 [0.8152, 0.8457]3.6575 [3.5037, 3.8412]2.8361 [2.7322, 2.9558]
Table 4. Execution time per random seed for each model.
Table 4. Execution time per random seed for each model.
ModelMean (s)Min (s)Max (s)
SVR0.01590.00200.0266
GBR0.05800.05370.0700
KNN0.06120.04440.0960
RF0.12040.10320.1766
ANN0.31920.23260.4352
XGBoost0.52430.31652.0587
HET0.65740.47112.1406
LightGBM0.67500.56920.8191
GPR1.59371.04742.3285
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Migallón, V.; Penadés, H.; Penadés, J. Key Predictors of Lightweight Aggregate Concrete Compressive Strength by Machine Learning from Density Parameters and Ultrasonic Pulse Velocity Testing. Mater. Proc. 2025, 26, 4. https://doi.org/10.3390/materproc2025026004

AMA Style

Migallón V, Penadés H, Penadés J. Key Predictors of Lightweight Aggregate Concrete Compressive Strength by Machine Learning from Density Parameters and Ultrasonic Pulse Velocity Testing. Materials Proceedings. 2025; 26(1):4. https://doi.org/10.3390/materproc2025026004

Chicago/Turabian Style

Migallón, Violeta, Héctor Penadés, and José Penadés. 2025. "Key Predictors of Lightweight Aggregate Concrete Compressive Strength by Machine Learning from Density Parameters and Ultrasonic Pulse Velocity Testing" Materials Proceedings 26, no. 1: 4. https://doi.org/10.3390/materproc2025026004

APA Style

Migallón, V., Penadés, H., & Penadés, J. (2025). Key Predictors of Lightweight Aggregate Concrete Compressive Strength by Machine Learning from Density Parameters and Ultrasonic Pulse Velocity Testing. Materials Proceedings, 26(1), 4. https://doi.org/10.3390/materproc2025026004

Article Metrics

Back to TopTop