Interpretable Ensemble Machine Learning for Liquefaction Risk Prediction
Abstract
1. Introduction
2. Materials
2.1. Dataset
2.2. Data Preprocessing
2.3. Classifiers
3. Methods
3.1. Dataset Correlation
3.2. Model Performance Metrics
3.3. Hyperparameter Optimization
3.4. Voting Ensemble
3.5. Feature Importance Analysis (SHAP)
4. Results and Discussion
4.1. Dataset Correlation
4.2. Base Models Performance
4.3. Voting Ensemble Performance
4.4. Comparison with Other Studies
4.5. Confusion Matrix
4.6. ROC and PR Curves
4.7. SHAP Analysis
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| Peak horizontal ground acceleration at the surface | |
| Uniform cyclic stress ratio | |
| Ground water table | |
| Moment magnitude of the earthquake | |
| Shear stress reduction coefficient (nonlinear shear mass participation factor) | |
| Total vertical overburden stress | |
| Effective vertical overburden stress | |
| Normalized shear wave velocity |
References
- Youd, T.L.; Idriss, I.M. Liquefaction Resistance of Soils: Summary Report from the 1996 NCEER and 1998 NCEER/NSF Workshops on Evaluation of Liquefaction Resistance of Soils. J. Geotech. Geoenviron. Eng. 2001, 127, 297–313. [Google Scholar] [CrossRef]
- Goh, A.T.C.; Goh, S.H. Support vector machines: Their use in geotechnical engineering as illustrated using seismic liquefaction data. Comput. Geotech. 2007, 34, 410–421. [Google Scholar] [CrossRef]
- Ishac, M.F.; Heidebrecht, A.C. Energy dissipation and seismic liquefaction in sands. Earthq. Eng. Struct. Dyn. 1982, 10, 59–68. [Google Scholar] [CrossRef]
- Seed, H.B.; Idriss, I.M. Analysis of Soil Liquefaction: Niigata Earthquake. J. Soil Mech. Found. Div. 1967, 93, 83–108. [Google Scholar] [CrossRef]
- Dobry, R.; Abdoun, T. Cyclic Shear Strain Needed for Liquefaction Triggering and Assessment of Overburden Pressure Factor Kσ. J. Geotech. Geoenvironmental Eng. 2015, 141, 04015047. [Google Scholar] [CrossRef]
- Robertson, P.K.; Woeller, D.J.; Finn, W.D.L. Seismic cone penetration test for evaluating liquefaction potential under cyclic loading. Can. Geotech. J. 1992, 29, 686–695. [Google Scholar] [CrossRef]
- Seed, H.B.; Idriss, I.M. Simplified Procedure for Evaluating Soil Liquefaction Potential. J. Soil Mech. Found. Div. 1971, 97, 1249–1273. [Google Scholar] [CrossRef]
- Choi, D.-H.; Kwon, T.-H.; Ko, K.-W. Centrifuge Modeling of Soil Liquefaction Triggering: 2017 Pohang Earthquake. KSCE J. Civ. Eng. 2024, 28, 3176–3191. [Google Scholar] [CrossRef]
- Ko, K.-W.; Lee, S.-B.; Oh, T.-S.; Kim, N.-R.; Kim, J.-H. Performance Evaluation of Sieve and Curtain Pluviators for Reconstituted Sand Specimen in Model Box. KSCE J. Civ. Eng. 2024, 28, 5452–5463. [Google Scholar] [CrossRef]
- Andrus, R.D.; Stokoe II, K.H. Liquefaction Resistance of Soils from Shear-Wave Velocity. J. Geotech. Geoenvironmental Eng. 2000, 126, 1015–1025. [Google Scholar] [CrossRef]
- Moon, S.W.; Mukhtarkhan, D.; Khamitov, R.; Abdialim, S.; Shokbarov, Y.; Kim, J.; Ku, T. Liquefaction Assessment using Surface Waves in Kazakhstan. In Proceedings of the 18th World Conference on Earthquake Engineering (WCEE2024), Milan, Italy, 30 June–5 July 2024. [Google Scholar]
- Goh, A.T.C. Seismic Liquefaction Potential Assessed by Neural Networks. J. Geotech. Eng. 1994, 120, 1467–1480. [Google Scholar] [CrossRef]
- Shao, W.; Yue, W.; Zhang, Y.; Zhou, T.; Zhang, Y.; Dang, Y.; Wang, H.; Feng, X.; Chao, Z. The Application of Machine Learning Techniques in Geotechnical Engineering: A Review and Comparison. Mathematics 2023, 11, 3976. [Google Scholar] [CrossRef]
- Fattah, E.A.; Ali, H.E.A.; Ebid, A.M. Prediction of Soil Liquefaction Using Genetic Programming. In Proceedings of the III Middel East Regional Conference on Civil Engineering Technology and III International Symposium on Environmental Hydrology, Dubai, United Arab Emirates, 29–30 September 2025. [Google Scholar]
- Pal, M. Support vector machines-based modelling of seismic liquefaction potential. Int. J. Numer. Anal. Methods Geomech. 2006, 30, 983–996. [Google Scholar] [CrossRef]
- Hu, J.-L.; Tang, X.-W.; Qiu, J.-N. Assessment of seismic liquefaction potential based on Bayesian network constructed from domain knowledge and history data. Soil Dyn. Earthq. Eng. 2016, 89, 49–60. [Google Scholar] [CrossRef]
- Hu, H.; Hu, X.; Gong, X. Predicting the strut forces of the steel supporting structure of deep excavation considering various factors by machine learning methods. Undergr. Space 2024, 18, 114–129. [Google Scholar] [CrossRef]
- Jas, K.; Dodagoudar, G.R. Liquefaction Potential Assessment of Soils Using Machine Learning Techniques: A State-of-the-Art Review from 1994–2021. Int. J. Geomech. 2023, 23, 03123002. [Google Scholar] [CrossRef]
- Das, S.K.; Mohanty, R.; Mohanty, M.; Mahamaya, M. Multi-objective feature selection (MOFS) algorithms for prediction of liquefaction susceptibility of soil based on in situ test methods. Nat. Hazards 2020, 103, 2371–2393. [Google Scholar] [CrossRef]
- Demir, S.; Sahin, E.K. Comparison of tree-based machine learning algorithms for predicting liquefaction potential using canonical correlation forest, rotation forest, and random forest based on CPT data. Soil Dyn. Earthq. Eng. 2022, 154, 107130. [Google Scholar] [CrossRef]
- Ozsagir, M.; Erden, C.; Bol, E.; Sert, S.; Özocak, A. Machine learning approaches for prediction of fine-grained soils liquefaction. Comput. Geotech. 2022, 152, 105014. [Google Scholar] [CrossRef]
- Kurnaz, T.F.; Erden, C.; Kökçam, A.H.; Dağdeviren, U.; Demir, A.S. A hyper parameterized artificial neural network approach for prediction of the factor of safety against liquefaction. Eng. Geol. 2023, 319, 107109. [Google Scholar] [CrossRef]
- Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 2–6 December 2024; pp. 4768–4777. [Google Scholar]
- Kayen, R.; Moss, R.E.S.; Thompson, E.M.; Seed, R.B.; Cetin, K.O.; Kiureghian, A.D.; Tanaka, Y.; Tokimatsu, K. Shear-Wave Velocity–Based Probabilistic and Deterministic Assessment of Seismic Soil Liquefaction Potential. J. Geotech. Geoenvironmental Eng. 2013, 139, 407–419. [Google Scholar] [CrossRef]
- Kohavi, R. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Proceedings of the 14th International Joint Conference on Artificial Intelligence—Volume 2, Montreal, QC, Canada, 20–25 August 1995; pp. 1137–1143. [Google Scholar]
- Kleinbaum, D.G.; Klein, M. Logistic Regression, 3rd ed.; Springer: New York, NY, USA, 2010; p. 702. [Google Scholar] [CrossRef]
- Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
- Cover, T.M.; Hart, P.E. Nearest Neighbor Pattern Classification. IEEE Trans. Inf. Theory 1967, 13, 21–27. [Google Scholar] [CrossRef]
- Friedman, N.; Geiger, D.; Goldszmidt, M. Bayesian Network Classifiers. Mach. Learn. 1997, 29, 131–163. [Google Scholar] [CrossRef]
- Hastie, T.; Tibshirani, R.; Friedman, J. The Elements of Statistical Learning, 2nd ed.; Springer: New York, NY, USA, 2009. [Google Scholar] [CrossRef]
- Breiman, L.; Friedman, J.H.; Olshen, R.A.; Stone, C.J. Classification and Regression Trees, 1st ed.; Chapman and Hall/CRC: Boca Raton, FL, USA, 1984; p. 368. [Google Scholar] [CrossRef]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
- Breiman, L. Bagging predictors. Mach. Learn. 1996, 24, 123–140. [Google Scholar] [CrossRef]
- Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef]
- Varma, S.; Simon, R. Bias in error estimation when using cross-validation for model selection. BMC Bioinform. 2006, 7, 91. [Google Scholar] [CrossRef] [PubMed]
- Yang, L.; Shami, A. On hyperparameter optimization of machine learning algorithms: Theory and practice. Neurocomputing 2020, 415, 295–316. [Google Scholar] [CrossRef]
- Feurer, M.; Klein, A.; Eggensperger, K.; Springenberg, J.T.; Blum, M.; Hutter, F. Learning. In Automated Machine Learning: Methods, Systems, Challenges; Hutter, F., Kotthoff, L., Vanschoren, J., Eds.; Springer International Publishing: Cham, Switzerland, 2019; pp. 113–134. [Google Scholar] [CrossRef]
- Şehmusoğlu, E.H.; Kurnaz, T.F.; Erden, C. Estimation of soil liquefaction using artificial intelligence techniques: An extended comparison between machine and deep learning approaches. Environ. Earth Sci. 2025, 84, 130. [Google Scholar] [CrossRef]
- Kumar, D.R.; Samui, P.; Burman, A.; Wipulanusat, W.; Keawsawasvong, S. Liquefaction susceptibility using machine learning based on SPT data. Intell. Syst. Appl. 2023, 20, 200281. [Google Scholar] [CrossRef]
- Jas, K.; Mangalathu, S.; Dodagoudar, G.R. Evaluation and analysis of liquefaction potential of gravelly soils using explainable probabilistic machine learning model. Comput. Geotech. 2024, 167, 106051. [Google Scholar] [CrossRef]
- Kohestani, V.R.; Hassanlourad, M.; Ardakani, A. Evaluation of liquefaction potential based on CPT data using random forest. Nat. Hazards 2015, 79, 1079–1089. [Google Scholar] [CrossRef]










| Classifiers | Descriptions |
|---|---|
| Logistic Regression (LR) | linear classification algorithm; predicts the probability of an instance belonging to a certain class using the logistic (sigmoid) function; operates under the assumption that the log-odds of the target variable being in a particular class is a linear combination of the input features [26]. |
| Support Vector Classifier (SVC) | part of the Support Vector Machine family; can be linear or nonlinear depending on the kernel parameter; seeks to find the optimal hyperplane that best separates different classes in the input data by maximizing the margin between classes [27]. |
| K-Nearest Neighbors (KNN) | non-parametric classification algorithm; classifies instances based on the majority class of nearest neighbors in the feature space; uses a distance metric such as Euclidean distance [28]. |
| Gaussian Naive Bayes (GNB) | probabilistic classifier based on Bayes’ theorem; assumes features are independent and normally distributed; simplifies the calculation process and is suitable for features with continuous values [29]. |
| Ridge Classifier (RC) | linear classification algorithm; incorporates L2 regularization to penalize large coefficients and prevent overfitting; useful for handling multicollinearity and improving model generalization [30]. |
| Decision Tree Classifier (DT) | non-parametric, tree-based classification algorithm; recursively splits the input space into regions based on feature values, selecting the feature that best separates the data; maximizes information gain or minimizes impurity measures like Gini impurity or entropy [31]. |
| Random Forest Classifier (RF) | ensemble learning algorithm with decision trees; combines multiple decision trees to improve accuracy and reduce overfitting; builds a forest of decision trees by training each tree on a random subset of features and instances, making predictions by aggregating the individual trees’ predictions [32]. |
| Bagging Classifier (BC) | ensemble learning algorithm with decision trees; builds multiple base classifiers using bootstrapped samples of the training data, combining their predictions through averaging or voting to reduce variance and improve accuracy [33]; by default, the base estimator for the BC in the “scikit-learn” library is a DT classifier. |
| Gradient Boosting Classifier (GB) | ensemble learning algorithm with decision trees; combines boosting and gradient descent techniques builds decision trees subsequentially, with each tree correcting errors of previous ones, optimizing a differentiable loss function using gradient descent, and aggregating predictions weighted by a learning rate [34]. |
| Classifier | Hyperparameter | Search Space |
|---|---|---|
| LR | penalty | L1, L2 |
| C | 0.1, 1, 10, 100 | |
| SVC | kernel | linear, rbf |
| C | 0.1, 0.2, 0.5, 1, 2, 5, 10, 100, 1000 | |
| gamma | 0.001, 0.002, 0.005, 0.01, 0.02, 0.05, 0.1, 1, 10 | |
| KNN | n_neighbors | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 |
| weights | uniform, distance | |
| algorithm | auto, ball_tree, kd_tree, brute | |
| GNB | var_smoothing | 1 × 10−6, 1 × 10−7, 1 × 10−8, 1 × 10−9, 1 × 10−10 |
| RC | alpha | 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0 |
| tol | 1 × 10−3, 1 × 10−4, 1 × 10−5 | |
| DT | criterion | gini, entropy, log_loss |
| max_depth | 3, 5, 7, 10, 20, 50, 100 | |
| RF | max_features | 3, 5, 7, 10, 20, 50 |
| max_depth | 3, 5, 7, 10, 20, 50, 100 | |
| BC | max_features | 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0 |
| max_samples | 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0 | |
| GB | learning_rate | 0.01, 0.1, 0.2, 0.5 |
| max_depth | 3, 5, 8 | |
| subsample | 0.8, 0.9, 0.95, 1 |
| Model | Accuracy (±stdv) | Precision (±stdv) | Recall (±stdv) | F1-Score (±stdv) | ||||
|---|---|---|---|---|---|---|---|---|
| LR | 82.65% | ±3.20% | 84.36% | ±2.26% | 91.98% | ±2.38% | 88.00% | ±2.22% |
| SVC | 86.75% | ±2.95% | 85.20% | ±2.42% | 97.91% | ±1.30% | 91.10% | ±1.87% |
| KNN | 81.69% | ±2.79% | 81.16% | ±2.18% | 95.83% | ±2.58% | 87.86% | ±1.86% |
| GNB | 78.31% | ±5.00% | 83.96% | ±4.16% | 85.00% | ±3.84% | 84.43% | ±3.55% |
| RC | 82.89% | ±3.18% | 82.93% | ±1.84% | 94.75% | ±3.51% | 88.43% | ±2.26% |
| DT | 86.02% | ±3.62% | 89.80% | ±4.35% | 90.25% | ±1.76% | 89.97% | ±2.46% |
| RF | 87.23% | ±3.54% | 87.53% | ±2.55% | 95.14% | ±2.27% | 91.17% | ±2.35% |
| BC | 89.16% | ±5.70% | 91.16% | ±3.93% | 93.39% | ±4.28% | 92.25% | ±4.05% |
| GB | 88.43% | ±4.97% | 88.63% | ±2.52% | 95.49% | ±4.95% | 91.91% | ±3.58% |
| Model | Accuracy (±stdv) | Precision (±stdv) | Recall (±stdv) | F1-Score (±stdv) | ||||
|---|---|---|---|---|---|---|---|---|
| LR | 83.61% | ±3.54% | 83.70% | ±2.13% | 94.75% | ±4.16% | 88.85% | ±2.61% |
| SVC | 85.78% | ±3.99% | 86.32% | ±2.58% | 94.42% | ±3.88% | 90.16% | ±2.79% |
| KNN | 85.30% | ±4.34% | 84.86% | ±4.56% | 96.18% | ±1.68% | 90.11% | ±2.77% |
| GNB | 78.31% | ±5.00% | 83.96% | ±4.16% | 85.00% | ±3.84% | 84.43% | ±3.55% |
| RC | 82.41% | ±3.62% | 83.01% | ±1.76% | 93.72% | ±4.06% | 88.02% | ±2.59% |
| DT | 85.30% | ±4.78% | 86.60% | ±5.53% | 93.74% | ±2.57% | 89.91% | ±2.96% |
| RF | 87.71% | ±5.67% | 87.68% | ±4.26% | 95.82% | ±3.57% | 91.55% | ±3.78% |
| BC | 87.47% | ±5.98% | 89.43% | ±4.59% | 93.02% | ±4.53% | 91.15% | ±4.10% |
| GB | 88.92% | ±5.46% | 89.48% | ±3.17% | 95.14% | ±4.67% | 92.22% | ±3.86% |
| Fold | Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|---|
| 1 | 81.93% | 84.13% | 91.38% | 87.60% |
| 2 | 91.57% | 90.48% | 98.28% | 94.21% |
| 3 | 90.36% | 90.16% | 96.49% | 93.22% |
| 4 | 95.18% | 93.44% | 100.00% | 96.61% |
| 5 | 91.57% | 89.06% | 100.00% | 94.21% |
| Mean | 90.12% | 89.45% | 97.23% | 93.17% |
| Stdv | 4.53% | 3.20% | 3.20% | 3.09% |
| Research Paper | Dataset (liq.:non-liq.) | Dataset Type | # of Inputs | Algorithm | Accuracy | Recall |
|---|---|---|---|---|---|---|
| This study | 415 (287:128) | 15 | Voting Ensemble | 90.12% | 97.23% | |
| [40] | 296 (159:137) | DPT | 5 | LightGBM | 83.33% | 83.33% |
| [21] | 273 (87:186) | CPT | 7 | DT | 89% | 77% |
| RF | 77% | 80% | ||||
| [20] | 226 (133:93) | CPT | 7 | CCF | 96.42% | 96.43% |
| [19] | 411 (287:124) | 9 | ANN + NSGA-II | 90.51% | 89.55% | |
| ANN + MOSOS | 92.94% | 92.68% | ||||
| [41] | 226 (133:93) | CPT | 6 | RF | 99.48% | N/A |
| [2] | 226 (133:93) | CPT | 6 | SVM | 97% | N/A |
| [12] | 85 (42:43) | SPT | 8 | ANN | 92% | N/A |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Tuzelbayev, D.; Moon, S.-W.; Lee, M.; Abdialim, S.; Aremu, E.A.; Satyanaga, A.; Kim, J. Interpretable Ensemble Machine Learning for Liquefaction Risk Prediction. Infrastructures 2025, 10, 304. https://doi.org/10.3390/infrastructures10110304
Tuzelbayev D, Moon S-W, Lee M, Abdialim S, Aremu EA, Satyanaga A, Kim J. Interpretable Ensemble Machine Learning for Liquefaction Risk Prediction. Infrastructures. 2025; 10(11):304. https://doi.org/10.3390/infrastructures10110304
Chicago/Turabian StyleTuzelbayev, Doszhan, Sung-Woo Moon, Minho Lee, Shynggys Abdialim, Elijah Adebayonle Aremu, Alfrendo Satyanaga, and Jong Kim. 2025. "Interpretable Ensemble Machine Learning for Liquefaction Risk Prediction" Infrastructures 10, no. 11: 304. https://doi.org/10.3390/infrastructures10110304
APA StyleTuzelbayev, D., Moon, S.-W., Lee, M., Abdialim, S., Aremu, E. A., Satyanaga, A., & Kim, J. (2025). Interpretable Ensemble Machine Learning for Liquefaction Risk Prediction. Infrastructures, 10(11), 304. https://doi.org/10.3390/infrastructures10110304

