Next Article in Journal
Exploring the AHP-AgileITS-ArchDesign: An AHP Model and Tool for Evaluating IT Service Architectural Agile Designs in SMBs
Previous Article in Journal
Co-Evolution of Artificial Intelligence and Green Technological Innovation: A Computational Mapping and Diagnostic Framework
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Machine Learning-Based Crisis Detection Framework for Banking Systems: A Case Study of Nigeria

by
Ntanganedzeni Mandiwana
1,
Thakhani Ravele
1,*,
Caston Sigauke
1 and
Rendani Netshikweta
2
1
Department of Mathematical and Computational Sciences, University of Venda, Private Bag X5050, Thohoyandou 0950, Limpopo, South Africa
2
Department of Mathematics and Applied Mathematics, University of Limpopo, Private Bag X1106, Sovenga 0727, Limpopo, South Africa
*
Author to whom correspondence should be addressed.
Analytics 2026, 5(3), 28; https://doi.org/10.3390/analytics5030028
Submission received: 15 May 2026 / Revised: 7 July 2026 / Accepted: 28 July 2026 / Published: 7 August 2026

Abstract

Banking crises are a persistent threat to macroeconomic stability in emerging markets, where conventional econometric monitoring frameworks often fail to capture non-linear macro-financial relationships. This paper examines whether machine learning algorithms can improve the detection of banking crisis risk in Nigeria compared to standard logistic regression. We compare the performance of Random Forest, Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost) against logistic regression using annual data from the African Financial Crises dataset (1954–2014). Resampling is only implemented on the training set to overcome the infrequency of crisis events. Performance on models is assessed based on accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUC) in a rigorous out-of-time validation setting. Our findings indicate that tree-based ensemble models outperform logistic regression on the test set: XGBoost achieves the best generalization performance (AUC = 1.0; F1 = 0.95 in non-crisis, 0.80 in crisis), whereas Random Forest has the highest cross-validated F1-score on the training set. The most important variables are exchange rate volatility, inflation, and indicators of systemic crisis. The most significant crisis indicators are, however, seen in crisis years, which means that the annual data do not provide much lead-time to detect the crisis. These results should be taken with caution because of the small sample size and the limited number of crisis observations during the test period. Altogether, machine learning models have potential as additional tools to monitor banking crises in Nigeria, though at the moment they are not fully operational as policy instruments.

1. Introduction

1.1. Overview

In emerging economies like Nigeria, where the banking industry plays a crucial role in both financial intermediation and economic development, tracking systemic vulnerabilities is a vital component of financial stability analysis. The Nigerian banking system experienced significant challenges during the early 1990s, the late 1990s, and  the global financial crisis (2009–2014), which resulted in structural weaknesses, such as high concentrations of non-performing loans, regulatory laxity, and  high sensitivity to macroeconomic and external shocks, [1,2]. These repeated crises highlight the need to have effective risk-monitoring systems that would help in detecting macro-financial vulnerabilities before they can result in systemic disruptions.
Traditionally, econometric models such as logit and probit regression models have been employed in monitoring frameworks for banking crises. These models are widely favored by financial regulators and policymakers due to their simplicity and clarity [3]. Nonetheless, they are limited by the high levels of linearity and the failure of the models to reflect the complicated relationships between the macro-financial variables. These constraints can lead to high false-positive rates and lower accuracy in emerging economies such as Nigeria, where the macroeconomic dynamics are characterized by non-linear relationships, and data constraints are severe, thus making traditional methods inapplicable in practice [3].
In recent years, machine learning (ML) has provided new methodologies that are particularly well adapted for extracting complex interactions and nonlinear relationships from high-dimensional financial signals, including gradient boosting machines [4]. The extensive empirical literature in this area shows that machine learning models, such as Random Forest, Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost), offer much greater precision in detecting macro-financial anomalies at the local level than traditional econometric specifications. This improved detection capability has been confirmed in both developed and leading emerging markets, including China and India [5,6,7]. In particular, ensemble learning approaches have shown strong promise in identifying subtle signals of systemic risk from macro-financial indicators, even under conditions of high economic volatility [6].
Although the global body of literature on ML-based crisis detection is expanding, the gap in the empirical research on the practice of applying the frameworks of ML to the banking industry in Nigeria is critical. This is especially worrying because Nigeria has a track record of macro-financial instability and requires efficient risk monitoring and systemic surveillance instruments [2]. To fill this gap, this paper compares the three ML models of the Random Forest, SVM, and XGBoost against a baseline logistic regression model, using historical data from the African Financial Crises database [8]. These algorithms are also tested on an out-of-time test partition, and their classification performance is measured using various metrics (accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (ROC-AUC)).
This research adds value to both the literature and the macroprudential policy-making process as it presents empirical data on trade-offs between model interpretability and classification accuracy in the case of the Nigerian context. The results offer data-driven conclusions to supplement existing central bank, financial regulator, and macroprudential authorities’ tracking tools by comparing non-parametric machine learning models with a conventional econometric baseline. Overall, these instruments contribute to prompt regulatory action to address emerging vulnerabilities, strengthen the financial stability pillars of the banking sector, and safeguard Nigeria’s economy.

1.2. Literature Review

1.2.1. Traditional Crisis Detection and Tracking Frameworks for Banking Crises

Crisis detection and tracking frameworks are created to identify weaknesses in the banking system so that immediate policy measures can be taken. In econometric models such as probit and logit regression, macro-financial variables are mostly employed in conventional monitoring models, including debt service ratios, exchange rate volatility, and credit-to-GDP ratios, among others [9]. The simplicity and ease of interpretation of these models make them very appealing to financial regulators and policymakers.
Nonetheless, there are some significant weaknesses of traditional monitoring models. The non-linear nature of financial regimes cannot be represented in their linear form, and this often leads to large false-positive rates, which negate their usefulness in practice. Such constraints are added in Nigeria because the economy is dependent on oil revenues, is prone to external shocks, and has difficulties accessing data [2].
Recent research highlights the need for more flexible crisis monitoring frameworks. Ref. [10] emphasizes the importance of crisis tracking indicators being consistent with policymakers’ decision-making cycles, while [11] demonstrates that advanced statistical and computational techniques can improve classification accuracy. The Nigerian banking crisis of 2009, caused by risk-taking, commodity price shocks, and exposure to stock markets, is an illustration of how detection tools are required that can both highlight country-specific vulnerabilities and global financial connections.
The vulnerabilities of the Nigerian banking system to structures further highlight the importance of advanced ML methods. In Nigeria, fiscal health is closely connected to financial stability, which is largely involved in unstable exogenous variables, especially the global oil price shocks that have contributed to the increase in non-performing loans in commercial banks [12,13,14]. Additionally, the industry has faced unpredictable regulatory measures, including the Central Bank of Nigeria’s aggressive consolidation policy in 2004, which compelled banks to merge rapidly [15,16], alongside persistent foreign exchange liquidity challenges and macroeconomic uncertainty [17,18]. The conventional linear tracking models simply cannot explain these non-linear abrupt changes and structural discontinuities. In contrast, ML models, with their flexible, high-dimensional structure, can capture these complex dynamics more accurately and reliably, offering a more comprehensive detection framework for emerging economies like Nigeria.
In order to explain the terminology of this study, we differentiate three concepts: crisis detection is the ability to detect vulnerabilities as they become real; crisis monitoring is the ongoing monitoring of macro-financial indicators over time; and forecasting/prediction is the ability to predict future risks with an adequate lead time (e.g., 12–24 months) to pre-emptive policy action. This paper focuses on crisis detection and monitoring, not forecasting or early warning, due to the annual frequency of the data and the empirical evidence of limited lead-time capacity (see Section 3.5).

1.2.2. Machine Learning Methods for Banking Crisis Detection

Machine learning (ML) presents a promising new approach to monitoring banking crises, as opposed to traditional econometric models. Unlike linear regression, ML methods can capture complex, non-linear relationships and interactions in high-dimensional data [19]. This flexibility renders ML especially well-suited to crisis detection frameworks of volatile and fragile financial systems.
Random Forest, SVM, and XGBoost have demonstrated improved classification accuracy over traditional methods [20,21]. Although logistic regression can be a good base because it is interpretable, ensemble and kernel-based ML algorithms can provide important improvements in accuracy and strength, especially in times of economic stress.
ML-based systemic risk monitoring frameworks in emerging economies have significant payoffs [1]. Nevertheless, oil shocks, like changes in oil prices, may lead to changes in financial stability in oil-exporting countries, including Nigeria [12]. Ensemble techniques such as Random Forest and XGBoost are better equipped to identify non-linear correlations among macro-financial variables compared to classical models [22].
Although they are more effective, the black-box character of the ML models can make them less appealing to policymakers who need their inputs to be explained to make policies [23]. Explainable AI (XAI) techniques, such as SHAP, have been proposed to improve model transparency by estimating the contribution of individual features. While XAI methods are valuable, this study focuses on assessing the ability of ML models to detect contemporaneous indicators of financial crises in Nigeria. It is recognized that XAI is a promising direction that can be used in future research.
Although SHAP and similar XAI methods are effective in other settings, this study aims to assess the ability of machine learning models to detect contemporaneous indicators of financial crises in Nigeria. Even though this falls out of the scope of the research, the use of explainable AI tools is also essential to discuss. This limitation gives subsequent research a chance to improve the classification performance and interpretation.
Although ML has demonstrated effectiveness in crisis detection globally, its application in Nigeria remains limited. The majority of studies centered on Nigeria are based on conventional econometric methods, which are characterized by large false-positive rates and are not able to reveal non-linear dynamics. Notably, no one has compared logistic regression and ML models in the Nigerian context in a systematic manner in order to establish the best trade-off between interpretability and classification accuracy. This empirical gap deprives policymakers of a clear direction on the choice of a model to detect a crisis.
To fill this gap, this study compares the performance of Random Forest, SVM, and XGBoost with logistic regression in terms of classification performance and its practical implications in macroprudential surveillance with Nigerian data. This study contributes to the literature by assessing crisis-tracking performance and acknowledging that the application of XAI approaches remains a topic for future research.
Table 1 in the companion source (literature comparison) gives a summary of the literature on how banking crises are tracked in the literature and how the methodology has been changing over time through the transition from traditional econometric models to machine learning algorithms, and how it may be applied to new markets like Nigeria.
This paper models the Nigerian banking crisis using logistic regression, Random Forest, Support Vector Machine, and XGBoost.

1.3. Contributions and Research Highlights

This study adds to the body of literature on financial stability and systemic risk monitoring by systematically evaluating the accuracy of the classification by the Random Forest, XGBoost, and SVM when compared to the logistic regression technique used in the detection of banking crises in Nigeria. The main contribution is demonstrating that ML-based methods, particularly XGBoost, achieve superior macroprudential tracking performance and offer a viable framework for crisis detection in developing countries.
The key highlights of this study are as follows.
  • Three ML models, Random Forest, XGBoost, and SVM, were evaluated against logistic regression for detecting financial crises in Nigeria.
  • Random Forest achieved the best cross-validation performance on the training data, with high F1-scores, recall, and ROC-AUC values.
  • XGBoost demonstrated the best generalization on unseen data, achieving the highest accuracy and crisis-class F1-score on the test set.
  • Exchange rate volatility, inflation, and systemic crisis indicators were found to be the most significant macro-financial variables.
  • A crisis tracking framework was created based on estimated crisis probabilities and was able to identify all historical crises in the red alert zone properly.
  • The research can offer policymakers empirical information to inform model choice and risk management to detect crises in Nigeria.
This research aims to determine whether advanced ML techniques outperform traditional logistic regression in tracking banking crises in Nigeria. Given the annual frequency of the dataset, the study adopts an exploratory benchmarking approach and acknowledges the limitations of the available data, including the small number of crisis events. The methodology involves feature selection and data preprocessing, followed by the application of multiple ML algorithms (SVM, XGBoost, Random Forest, and Logistic Regression) to capture potential non-linear relationships. Since the dataset has an imbalance in classes, we measure model performance in terms of F1-score and ROC-AUC, as these are the right measures to use when dealing with an imbalanced classification problem.
The remainder of this paper is structured as follows. Section 2 outlines the modeling framework. Section 3 presents the empirical findings. Section 4 discusses the results, and Section 5 concludes.

2. Methods

In this section, the methodological framework will be described to assess whether machine-learning-based crisis detection frameworks are effective in detecting banking crises in Nigeria. The method is meant to be transparent, repeatable, and robust, as line with best practices for monitoring systemic risks. Its methodology includes the description and preprocessing of the data, the choice of features, model training based on both conventional econometric and machine learning approaches, and evaluation of the model based on cross-validation and out-of-time testing.

2.1. Data Description and Preprocessing

The macro-financial data of Nigeria is analyzed with the African Financial Crises Dataset, covering the 1954–2014 period, which consists of about 60 annual observations [8].
The variables included in the dataset are exchange rate volatility (calculated as the standard deviation of yearly changes in the exchange rates), inflation (year-over-year CPI), credit-to-GDP ratio, indicators of sovereign default in external debt, a systemic crisis dummy variable that represents concurrent crises in other sectors, debt service ratios,  oil price exposure variables, and current account balance variables.
Banking crises are rare in the dataset, with only approximately 11 crisis years. As a result, data preparation was done with caution to avoid overfitting and information leakage. Less than 5% of the data were missing; continuous variables were imputed using the mean, while binary variables were imputed using the mode. The continuous variables were mean-centered and variance-centered to zero and one, respectively, to improve numerical stability and make the models more comparable.
Given the infrequent occurrence of banking crises relative to non-crisis periods, the data is highly imbalanced. To overcome this, the Synthetic Minority Over-sampling Technique (SMOTE) was only used on the training set and not on the test set. The data were temporally split into a training sample (1954–2000) and a distinct out-of-time test sample (2001–2014). This method is consistent with the practice of risk tracking in the real world and does not introduce look-ahead bias. It also ensures that synthetic data do not influence the evaluation phase and that information leakage from the training data to the test period is prevented. The minority class is used for crisis years, and resampling was used to help models more effectively identify crisis years without changing the initial class distribution during testing.
Several precautions were implemented to avoid bias or distortion of the feature space in the synthetic data. To begin with, SMOTE was implemented only after splitting by temporal scale so that the test set (2001–2014) was left intact and a true reflection of the banking situation in Nigeria. Second, stratified cross-validation was applied to make sure that synthetic instances did not saturate but augmented class boundaries. Evaluation metrics (ROC-AUC and Recall) on the out-of-time test sample confirm the validity of this approach, demonstrating the models’ ability to generalize to real-world crisis instances without being misled by synthetic noise.
A key constraint is the relatively small number of observations ( n 60 ) and the limited number of crisis years (≈11). This limits the complexity of the models that can be estimated with confidence. To prevent overfitting and unstable parameter estimates, the following strategies were implemented: (i) a parsimonious feature set selected using XGBoost importance filtering (see Section 2.2); (ii) the use of regularized estimators where appropriate; (iii) five-fold stratified cross-validation; (iv) outlier capping at the first and 99th percentiles; and (v) conservative interpretation of performance measures with explicit uncertainty quantification.

2.2. Feature Selection

Given the small number of observations and the relatively large number of macro-financial variables, feature selection was employed to improve model generalizability and robustness. A Gradient-Boosting-based feature selection method was employed because it can detect non-linear relationships and interaction effects, which are often found in macro-financial data.
The Gradient Boosting Classifier was trained on the SMOTE-resampled training data, and the importance of the features was estimated by a gain-based metric, which measures how much each feature contributes to the classification performance of the model. The variables whose importance scores were below 0.01 were dropped, whereas those whose scores were above the value were included based on their economic importance.
The original set of candidates consisted of 11 indicators such as case, year, systemic_crisis, exch_usd, domestic debt in default, sovereign external debt default, gdp weighted default, inflation annual cpi, independence, currency crises, and inflation crises. Following the screening process, the retained variables were “exch_usd” (exchange rate volatility), “year” (temporal index), “systemic_crisis” (systemic crisis indicator), and “inflation_annual_cpi” (inflation). In order to have complete reproducibility, Table 2 provides the feature importance scores and selection status of all candidate variables.
The “year” variable was initially included in the exploratory candidate set to assess baseline chronological trends; however, we recognize that a sequential temporal index does not represent a structural macro-financial indicator. In order to overcome the possible temporal overfitting issues, a strict robustness test was performed by retraining and retesting all models without the year variable (see Section 3.7).
The selection of features is the process that balances statistical relevance and economic theory, which is crucial especially when the idea is to track something based on policy.

2.3. Model Development

This research uses four classification models that indicate various modeling strategies that can be used in crisis monitoring frameworks: a linear probabilistic model with regularization (Logistic Regression), an ensemble bagging model (Random Forest), a margin-based model with kernel transformation (Support Vector Machine), and an ensemble boosting model (Extreme Gradient Boosting). This diverse set enables comparison between traditional econometric approaches and machine learning techniques.
Given the small dataset and limited number of crisis observations, several measures were implemented to control model complexity and avoid overfitting:
(i)
Feature Selection: Feature importance analysis was performed using XGBoost to identify important variables and reduce the dimensionality of the models, focusing on the most relevant macro-financial indicators.
(ii)
Cross-Validation: To get good performance estimates and to avoid overfitting to the training splits, five-fold stratified cross-validation is used in the training procedure.
(iii)
Hyperparameter Tuning: For ensemble models, hyperparameter tuning was performed using grid search with cross-validation to choose hyperparameters that perform well on unseen data.
(iv)
Regularization: L2 regularization was applied to the Logistic Regression model (via scikit-learn’s default C = 1.0 with solver = ‘liblinear’) to penalize large coefficients.
(v)
Class Imbalance Handling: SMOTE was applied exclusively to the training folds, and class_weight = ‘balanced’ (or scale_pos_weight for XGBoost) was used to adjust class weights during training.
(vi)
Outlier Capping: There was a numeric capping of the extreme values (first and 99th percentiles) to reduce the effects of outliers on parameter estimates.
The models were trained on the preprocessed training data (1954–2000) and tested on a separate out-of-time period (2001–2014).

2.3.1. Logistic Regression

Logistic regression is used as a benchmark model because it is widely used to track systemic risk and is interpretable. This model is often favored by policymakers and regulators due to the fact that the estimated coefficients can be directly related to economic theory in order to provide clear communication and decision-making.
The model estimates the probability of a banking crisis ( y = 1 ) as
P ( y = 1 X ) = 1 1 + exp ( β 0 + β X ) ,
where X represents the set of macro-financial variables and β denotes the coefficients estimated using maximum likelihood.
L2 regularization is used by default (C = 1.0 and solver = ‘liblinear’) to prevent overfitting and to encourage small coefficients, which is important in light of the few observations of the crisis. The model is used as a baseline to test the view of whether ensemble methods offer further gains in classification performance.

2.3.2. Random Forest

Random Forest is an ensemble learning algorithm that constructs multiple decision trees on bootstrapped samples of the data using random subsets of features. This method minimizes variance by averaging with ensembles and overfitting, so it is suitable for small macro-financial data.
The algorithm finds applications especially in crisis detection since it can automatically derive non-linear relationships, threshold effects, and interactions without necessarily specifying functional forms. For a forest with K trees, classification is made by majority voting:
y ^ = arg max c k = 1 K I ( y ^ k = c ) ,
where I ( · ) is the indicator function. Hyperparameter tuning was done using GridSearchCV and five-fold stratified cross-validation on the training set. The search space was
n estimators { 50 , 100 , 200 } , max _ depth { 5 , 10 , 20 } , min _ samples _ split { 2 , 5 , 10 } ,
and
min _ samples _ leaf { 1 , 2 , 4 } .
The optimal parameters were chosen to minimize the cross-validated log loss. Moreover, class_weight = ‘balanced’ was applied to balance the classes in training.

2.3.3. Support Vector Machine

Support Vector Machines (SVMs) are margin-based classifiers that seek to identify the optimal decision boundary between crisis and non-crisis observations. SVMs are especially useful in high-dimensional, small-sample settings, which is evident in macro-financial tracking frameworks.
The maximization of the separation between the classes enhances generalization and resilience to noisy features, which is particularly valuable in situations when the number of crisis observations is low, and they may also be affected by measurement error. The optimization problem is formulated as
min 1 2 w 2 + C i ξ i ,
subject to y i ( w x i + b ) 1 ξ i and ξ i 0 . A radial basis function (RBF) kernel is employed to model nonlinear patterns:
k ( x , x ) = exp ( γ x x 2 ) .
Hyperparameter tuning was done using GridSearchCV with five-fold stratified cross-validation, and the hyperparameters searched were
C { 0.1 , 1 , 10 }
and
γ { 0.01 , 0.1 , 1 } .
The selected values ( C = 1.0 , γ = 0.1 ) balance model flexibility with robustness on the imbalanced dataset. class_weight=’balanced’ was used to mitigate the impact of class imbalance.

2.3.4. Extreme Gradient Boosting

Extreme Gradient Boosting (XGBoost) is an advanced ensemble learning algorithm that builds decision trees iteratively, with each new tree correcting errors from previous trees. This adaptive learning approach enables XGBoost to effectively identify systemic risk signals associated with rare but severe events such as banking crises.
XGBoost is effective for crisis detection due to its ability to learn complex non-linear relationships, handle imbalanced data, and employ regularization to prevent overfitting. These characteristics are especially useful in the modeling of macro-financial systems that are exposed to structural changes and regime shifts.
The objective function minimized by the model is
L = i l ( y i , f k ) k Ω ( f k ) ,
in which l ( · ) is the logistic loss and  Ω ( · ) is the penalty of model complexity. The algorithm takes advantage of the second-order gradient data to achieve effective convergence.
Hyperparameter tuning was done with GridSearchCV with five-fold stratified cross-validation, and the hyperparameters were tuned over
learning _ rate { 0.01 , 0.1 , 0.2 } , n estimators { 50 , 100 , 200 } , max _ depth { 3 , 5 , 7 } ,
and
subsample { 0.6 , 0.8 , 1.0 } .
To achieve the lowest cross-validated log loss, the final parameters were set to learning_rate = 0.1, n_estimators = 100, max_depth = 5, and subsample = 0.8, with class imbalance addressed by the use of scale_pos_weight.
The four models were chosen due to their inductive biases and applicability to small samples. L2 regularization gives a stable and interpretable baseline using Logistic Regression. Random Forest and XGBoost use ensemble averaging and shrinkage, respectively, to minimize variance and overfitting. High-dimensional, small-sample problems are theoretically well-posed with SVMs using RBF kernels due to their margin-maximization target. All performance estimates are accompanied by confidence intervals (see Section 4) to acknowledge the uncertainty associated with the limited number of crisis observations.

2.4. Model Evaluation and Interpretability

Accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (ROC-AUC) are used to evaluate model performance. Given the class imbalance and the policy-relevant need to correctly classify crisis periods, greater emphasis is placed on recall and F1-score for the crisis class. Accuracy is also deceptive in imbalanced environments where a model might be very accurate simply by being able to predict the times when there is no crisis. Another application of ROC-AUC is to assess the total discriminatory power at all classification levels.
Models were validated using five-fold stratified cross-validation on the training set and an out-of-time evaluation on the test period (2001–2014). Gain-based feature importance scores from the XGBoost model were used to enhance interpretability. Gain-based importance addresses the black-box problem by giving quantitative scores to features that are proportional to the overall decrease in the loss function, which can be attributed to splits on that feature in all trees. This method is used to determine the most important factors of crisis detection, including exchange rate volatility and inflation, and can be tested against economic theory.
Gain-based importance does not tell which variable influences the ensemble model the most, or which ones have the least impact; however, it can be used to obtain an idea of the most influential variables without necessarily having to explore the inner workings of the ensemble model. Any probability higher than 0.5 was treated as a red alert in the crisis tracking system.
The evaluation methodology favors transparency and replicability, especially because of the shortcomings of annual macro-financial data and the rarity of banking crises. Although the method allows systematic comparison of ML models based on high validation standards, the results can only be considered as a hint and not an absolute indication of the possible ability to track in real time.

3. Results

3.1. Exploratory Data Analysis

Before presenting the model estimation results, this section provides a brief description of the Nigerian banking crisis dataset.
The dataset contains 60 annual observations for Nigeria from 1954 to 2014; 11 of these years (18.3%) are classified as banking crisis years, while the remaining 49 years are non-crisis years. This asymmetry in crisis and non-crisis observations is a typical feature of financial crisis data and the rationale behind the application of resampling methods to model training.
Table 3 provides descriptive statistics for the full set of macroeconomic and crisis-related variables. The descriptive statistics indicate that there is a great difference in significant indicators. Specifically, exchange rate fluctuations are quite dispersed, with phases of steep currency depreciation under the influence of exogenous shocks and fluctuations in oil prices. The annual inflation (CPI) is also very volatile, with the annual inflation in excess of 70% in times of macroeconomic instability, implying that there is high economic volatility over the sample period.
Figure 1 presents a correlation heatmap of the macroeconomic indicators. The correlation between exchange rate volatility and inflation is positive, and periods of high inflationary pressure and exchange rate volatility are associated with sovereign external debt defaults. These preliminary correlations suggest that macroeconomic volatility is associated with banking crises in Nigeria.
A notable finding from the correlation analysis is the high positive correlation (0.94) between systemic_crisis and banking_crisis. This is expected, as banking crises are a subset of systemic crisis events in the dataset. To conduct the analysis within the framework of this study, which is aimed at detecting contemporary but not advanced crises, the indicator of systemic crises is conceptually relevant: the model is used to determine whether a systemic crisis event is a banking crisis. However, readers should take note of this overlap in the interpretation of later results, as it can be a source of the high classification performance seen.
Generally, the exploratory analysis confirms that the dataset captures economically significant variation and crisis-related information, supporting its appropriateness for systemic risk-tracking frameworks.

3.2. Feature Importance

Feature importance was analyzed using the Gradient Boosting and Random Forest algorithms to identify the most important determinants of banking crises in Nigeria. The results are presented in Figure 2.
The most important indicators were exchange rate volatility (importance = 0.425), systemic crisis (importance = 0.139), and inflation (importance = 0.131). These variables continuously proved to be the major forces behind detection of crises. All the other candidate variables had an importance score of zero and were thus not included in the final model; these include sovereign external debt default. The year variable was not tested but was added to facilitate exploratory results that would be later tested in a robustness check (see Section 3.7), which showed that this omission does not materially impact model performance.

3.3. Model Performance

3.3.1. Cross-Validation

The initial evaluation of model performance was with five-fold cross-validation of the SMOTE-resampled training dataset. Table 4 reports the average performance metrics.
Random Forest produced the best overall results with an F1-Score of 0.93, Recall of 0.95, and ROC-AUC of 0.99. XGBoost also demonstrated strong performance, whereas Logistic Regression exhibited moderate classification performance. In contrast, the Support Vector Machine (SVM) model performed relatively poorly.

3.3.2. Test Set Evaluation

While Random Forest achieved the strongest performance during cross-validation on the training data, XGBoost outperformed all other models on the strictly out-of-time test set. Test set performance metrics are reported in Table 5. XGBoost achieved the highest accuracy (0.92) and the strongest crisis-class F1-score (0.80). Random Forest and Logistic Regression also gave good results. The test set is an out-of-time test (2001–2014), which is only a realistic evaluation of explanatory capacity under current historical detection conditions.
However, there are some key caveats to be considered. The test set comprises only 12 annual observations, with twp crisis years (2012, 2013) and 10 non-crisis years (including 2002 and 2008 as illustrative examples). This extremely limited and homogeneous test sample severely constrains the reliability of our performance metrics. This means that the near-perfect ROC-AUC values (1.00) that are being presented in Table 5 need to be considered as a proof of concept but not as evidence of the generalizability of the model.
To quantify the uncertainty arising from this small test sample, we computed bootstrapped 95% confidence intervals (1000 iterations, sampling with replacement from the test set) for all performance metrics. These are presented as “Mean [95% CI]” in Table 5.
The [1.00, 1.00] confidence intervals arise because the test set is extremely small (12 observations). With so few observations, each bootstrapped resample contains the same limited set of crisis and non-crisis years, consistently yielding perfect AUC values. This leads to zero-variance intervals, which are indicative of test sample limitations and not indicative of model perfection. The near-perfect ROC-AUC values (1.00) reported in this study are largely attributable to the extremely small out-of-time test sample (12 observations). These findings are to be understood as exploratory and as a proof of concept, as opposed to evidence of generalizable predictive performance.
We observe that such test set results are not evidence of generalizable perfection but do suggest that the models are able to discriminate crisis years and non-crisis years within the confines of this historical period. These findings should be validated with larger hold-out samples in future work.

3.4. ROC and Precision–Recall Analysis

Figure 3 and Figure 4 show the model’s discrimination ability using Precision–Recall curves and ROC curves. XGBoost and Random Forest perform better than other models, as they can almost perfectly distinguish crisis and non-crisis data points.

3.5. Systemic Risk Tracking Framework Outputs

The risk monitoring framework was developed using the predicted crisis probabilities from the XGBoost model, which achieved the best generalization performance on the out-of-time test set. Risk levels were categorized into Green (low risk), Yellow (moderate risk), and Red (high risk). Figure 5 displays the crisis detection timeline, including estimated probabilities of crises as well as actual banking crises.
The 11 historical crises are all in the Red alert zone, which shows that there is a high level of crisis detection. However, predicted probabilities tend to increase significantly during crisis years themselves, with minimal lead time at t 1 or earlier. This supports the analysis that the framework is a modern detection device as opposed to a proactive early warning system. Though the system is very effective at capturing crises as they arise, it lacks good indicators that can be used to initiate the policy proactively.

3.6. Summary of Key Findings

The exploratory analysis provides evidence of the relative performance of machine learning models for banking crisis risk identification in Nigeria under severe data constraints. The key findings are summarized as follows:
  • Tree-based ensemble models, especially the Random Forest and XGBoost, are always more effective in detecting historical banking crisis events, indicating that they are able to capture complex macro-financial trends.
  • Exchange rate volatility, inflation and systemic crisis indicators are the most significant indicators across model specifications that illustrate the significance of macroeconomic instability in the Nigerian banking industry.
  • Given the annual frequency of the data, the model is best suited for contemporaneous crisis detection rather than forward-looking identification. Crisis probabilities are highest during crisis years themselves, indicating limited lead time. Nevertheless, all historical crisis episodes were correctly identified within the alert zone, demonstrating the framework’s effectiveness in identifying crisis years as they occur.
  • The findings are to be read as suggesting performance above average classification, not necessarily as showing the presence of a working detection mechanism, but the framework is correct in pointing to the presence of historical crisis cases.
  • In general, the results indicate that machine learning algorithms can be used to supplement traditional econometric techniques to monitor the risk of a banking crisis in Nigeria on a real-time basis. Nonetheless, the application of the system in practice will be subject to the conditions of more comprehensive data accessibility, more transparent lead-time performance, and better interpretability.
When interpreting these results, it is important to consider the limitations of the data and the evaluation framework. However, the findings show that tree-based ensemble algorithms, especially Random Forest and XGBoost, have a high accuracy in classification. The resampling techniques, annual data, and few cases of crisis indicate that performance outcomes are an indication of favorable learning conditions but not an assurance of the real-time detection of crises. The next section discusses these empirical results within the context of the literature, their interpretation, and their limitations for the development of risk detection frameworks in emerging markets.

3.7. Robustness Check: Exclusion of Temporal Index

As the “year” variable was found to be one of the most important variables (importance: 0.305) in feature selection, a robustness test was performed to determine whether the models were overly dependent on temporal indexing rather than substantive economic indicators. Feature selection and model training were repeated without the “year” variable in the candidate feature set.
Without ‘ ‘year”, the features with importance above the selection threshold (importance > 0.01 ) were “exch_usd” (exchange rate volatility, importance: 0.425), “systemic_crisis” (importance: 0.139), and “inflation_annual_cpi” (importance: 0.131). The remaining features had higher absolute importance scores because they were rescaled when “year” was omitted, allowing the analysis to focus solely on economic and crisis indicators.
Table 6 compares the test set performance of models trained with and without the “year” variable. The performance numbers were relatively similar across all the models. The most notable change was an increase in Random Forest’s accuracy from 0.75 to 0.83. Perfect ROC-AUC and PR-AUC scores of 1.00 were maintained by all models and perfect recall (1.00) was preserved for the crisis class.
The robustness test confirms that the models’ ability to detect banking crises is not solely determined by the “year” variable. The economic indicators (exch_usd, systemic_crisis, inflation_annual_cpi) are sufficiently informative to maintain high explanatory power. This enhances trust in the structure, which means it does not rely on time mysteries but rather on substantive economic indicators. When the model works well without the temporal index, it is more likely to be a good indication of actual underlying economic relationships and less likely to be overfitting to certain time periods. We note, however, that the small test set size ( n = 12 ) may have contributed to the perfect scores, and further validation on larger test sets is required.

4. Discussion

This study demonstrates that ensemble machine learning algorithms, particularly Random Forest and XGBoost, are more effective than traditional logistic regression in detecting past banking crises in Nigeria [24,25]. Random Forest demonstrated excellent results on in-sample data, with a recall of 0.95 and an F1-score of 0.93, which indicates its capability to learn the trends in historical crisis events via bagging and deep decision trees. In contrast, XGBoost generalized better on the out-of-time test dataset, achieving an accuracy of 0.92. This improved performance is likely attributable to its sequential boosting process and regularization techniques, which enhance the model’s ability to handle temporal uncertainty.
The analysis of the importance of features showed that exchange rate volatility, systemic crisis, and inflation are the most explanatory factors, which proves that macroeconomic instability is the decisive factor of financial vulnerability in emerging economies. The influence of the year variable can probably be explained by structural changes, such as financial cycles in the world and changes in policy.
It is important to acknowledge that systemic_crisis and banking_crisis exhibit a high correlation (0.94), reflecting that banking crises are a subset of systemic crisis events. This overlap has to be taken into account when the feature importance of systemic crisis is interpreted. Although the variable helps in the high detection ability of the model, the inclusion of the variable implies that the model is in part exploiting information conceptually related to the target variable. This does not invalidate the validity of the detection framework, but it does indicate that performance measures like ROC-AUC and F1-score can be overstated when compared to a model with only independent macroeconomic indicators. Future research ought to examine whether the same detection accuracy is attainable with only variables that are conceptually different from the crisis outcome.
While ensemble methods (Random Forest and XGBoost) achieved higher accuracy and F1-scores on the out-of-time test set compared to Logistic Regression, we caution against attributing this performance difference solely to the ability of ensemble methods to capture non-linear relationships. Bootstrap-based statistical comparisons of ROC-AUC performance revealed no statistically significant differences between linear and non-linear models (p-values close to 1.0000), with mean ROC-AUC differences of 0.0000. This lack of significance is likely due to the perfect ROC-AUC (1.00) achieved by all models on the very small test set (12 observations), which prevented statistical differentiation. Therefore, no direct empirical evidence of non-linear effects on performance differences can be established. To definitively determine whether performance benefits arise from non-linear relationships, formal non-linearity testing (e.g., Ramsey RESET) would need to be conducted on larger datasets.
The temporal validation framework (out-of-time testing) provides a useful template for assessing model robustness that can be replicated in other contexts. The primary challenge in applying this approach to datasets with different properties, such as higher-frequency data, cross-country panels, or more frequent crisis events, is appropriately handling the minority class of crisis observations. In other crisis modeling problems, methods such as SMOTE have been successfully combined with temporal cross-validation to address this issue.
However, the temporal validation framework has structural constraints that must be explicitly recognized. The out-of-time test partition (2001–2014) provides a genuine test interval but is based on only 12 annual observations with a very sparse distribution of banking crisis events. Therefore, the optimal ROC-AUC values (1.00) achieved within the models are not stable in the long term and should not be used as an indication of absolute generalization. Instead, they imply that the models can be useful in monitoring crises in a narrow band of macro-financial deterioration. These high discrimination levels require confirmation in future studies using larger and more flexible target horizons and multi-country panels. It is important to reiterate that the near-perfect ROC-AUC values (1.00) are a direct consequence of the extremely small test set (12 observations) and should not be interpreted as evidence of generalizable predictive performance. Instead, they suggest that the models are viable in the particular context of this historical time frame.
As anticipated in the literature, tree-based ensemble models tend to outperform linear econometric models in detecting emerging-market crises [24,25]. Nonetheless, other studies have observed that standard logistic regression can outperform machine learning models in recursive out-of-sample tests [20].
A granular performance analysis was performed to evaluate the classification mechanics in terms of the imbalance in the classes to evaluate the minority crisis class (Label 1). Although all models had a great baseline capacity to explain historical systemic shocks (test recall of 1.00), significant variations in precision were found. The false-alarm rate for traditional logistic regression was high, with test precision of only 0.50 (CV F1: 0.7569; CV Recall: 0.7393; CV ROC-AUC: 0.9638). The global linear loss function, which optimizes total loss over the entire data space, can distort the decision boundary due to the majority class, leading to a high number of false-positive errors.
Support Vector Machine (SVM) was a special case, with perfect recall of 1.00 and an out-of-time test ROC-AUC of 0.00. The diagnostic examination of the model posterior distributions indicated that there was an inversion of the ranking; the SVM gave an identical probability estimate of 0.5000 to genuine crisis years and slightly more (0.5100 to 0.5330) to years without a crisis. All non-crisis instances scored higher than crisis instances, leading to ideal threshold mapping. The order of the scale ( 1 P ) shifted the ROC-AUC to 1.00, but the stability of cross-validation was low (CV F1: 0.4682; CV Recall: 0.5000; CV ROC-AUC: 0.8821).
Conversely, Random Forest had a high resampled training space dominance (CV F1: 0.9270; CV Recall: 0.9500; CV ROC-AUC: 0.9953), but with a test precision of 0.40. XGBoost, the most effective framework, had the highest test precision of 0.67 and an out-of-time F1-score of 0.80 (CV F1: 0.8687; CV Recall: 0.8679; CV ROC-AUC: 0.9828) and balanced best between target discrimination and false-alarm mitigation.
Though such ensemble architectures have obvious detection advantages, the empirical lead time of the monitoring system is limited by the underlying data, especially its frequency per year. The acute short-term liquidity shocks or abrupt currency runs that may occur mid-year may lead to a systemic banking crisis. Annual aggregation of macro-financial aggregates smooths and masks these rapid inter-annual changes. Visual inspection of the alert timelines confirms that crisis probabilities are very high either in the crisis year itself or, at most, in the preceding year. This framework is thus more of an accurate tracking system or immediate looming sensor rather than a long-term monitoring system. Further studies ought to overcome this limitation by using higher-frequency proxy measures (monthly or quarterly), incorporating more institutional and global risk measures, and testing more sophisticated sequential architectures like Recurrent Neural Networks (RNNs) or Long Short-Term Memory (LSTM) networks, along with explainable AI tools like SHAP, to elucidate current relationships.
Besides the limitations imposed by the data and time, there are other restrictions that should be mentioned. First, the test set comprises only 12 annual observations, with two crisis years (2012, 2013) and 10 non-crisis years. This very limited and homogeneous test sample yields artificially small confidence intervals and practically perfect point estimates (AUC = 1.00 across all models). These results are to be considered as exploratory and proof-of-concept as opposed to evidence of model generalizability. Second, the findings are country-specific in Nigeria and cannot necessarily be extended to other emerging markets with varying institutional structures, policy regimes, or history of previous crises. Third, the importance of gain-based features cannot be interpreted comprehensively because of the complexity of the Random Forest and XGBoost in nature, which makes the models less appropriate in policy communication, which may need causal explanation. Lastly, the model stability over time can be influenced by structural breaks and changes in policy regimes in Nigeria throughout the 60-year sample period. Collectively, these limitations affect the interpretability and generalizability of our results and indicate the necessity to validate them on larger and multi-country datasets.

5. Conclusions

This study demonstrates that machine learning algorithms, particularly tree-based ensemble models, can be more effective at detecting banking crises in Nigeria than traditional logistic regression. The findings indicate that the most important factors associated with banking system distress are macroeconomic uncertainty, as measured by exchange rate volatility, inflation, and systemic crisis indicators. However, the results are constrained by the limited number of banking crises, and the models are unable to provide indicators with substantial advance notice. Instead, they identify increasing risk contemporaneously as crises materialize, making them suitable for monitoring purposes.
A key limitation is that the near-perfect ROC-AUC values (1.00) reported in this study arise from the extremely small out-of-time test sample (12 observations). Consequently, these results can be considered exploratory and as a proof of concept instead of as conclusive evidence of predictive performance that can be generalized. These findings ought to be confirmed in future studies using larger, more diverse sets of data.
A limitation of this study is the conceptual overlap between systemic_crisis and banking_crisis, which are highly correlated (0.94). While this overlap is justified for a contemporaneous detection task, where the goal is to identify whether a systemic crisis is a banking crisis, it may contribute to the strong performance metrics reported. Consequently, the results are to be viewed as an indication of the successful classification of the crisis but not as the ability to forecast on its own. Further study should aim to confirm these findings with only macroeconomic variables that are conceptually independent of the crisis outcome and should incorporate higher frequency data to enhance lead-time possibilities.
This paper shows that machine learning can be used with traditional econometric techniques in data-sparse contexts. It offers a well-defined assessment framework that puts into perspective the potential as well as the limitations of such approaches. The models are to be viewed cautiously and are not to be taken as operational early warning systems as they are. Future research should focus on the use of high-frequency proxy data and multi-country panel models to improve classification robustness and macroprudential policy relevance.

Author Contributions

Conceptualization, N.M., T.R. and R.N.; methodology, N.M.; software, N.M.; validation, N.M., R.N., T.R. and C.S.; formal analysis, N.M.; investigation, N.M., R.N. and T.R.; data curation, N.M.; writing—original draft preparation, N.M.; writing—review and editing, N.M., R.N., T.R. and C.S.; visualization, N.M.; supervision, T.R., R.N. and C.S.; project administration, T.R. and R.N.; funding acquisition, N.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the DST-CSIR National e-Science Postgraduate Teaching and Training Platform (NEPTTP): http://www.escience.ac.za/ (accessed on 1 January 2025).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data were obtained from the Kaggle website https://www.kaggle.com/jaimetrickz/african-crises (accessed on 16 May 2025) [26]. The analytic dataset used in this study was derived by filtering and processing the original Nigerian data and is available from the corresponding author upon reasonable request.

Acknowledgments

The authors gratefully thank the support given to this research by the DST-CSIR National e-Science Postgraduate Teaching and Training Platform (NEPTTP). All opinions, findings, and conclusions expressed are the sole responsibility of the authors and are not necessarily those of NEPTTP. The authors also thank the anonymous reviewers for their constructive and insightful feedback.

Conflicts of Interest

In this manuscript, the authors stated that there are no conflicts of interest. The funding bodies were not involved in the design of the study, data collection, analysis, interpretation, manuscript writing, or the decision to publish findings.

Abbreviations

The abbreviations listed in this manuscript are utilized:
MLMachine Learning
SVMSupport Vector Machine
RFRandom Forest
XGBoostExtreme Gradient Boosting
GDPGross Domestic Product
CPIConsumer Price Index
ROCReceiver Operating Characteristic
AUCArea Under the Curve
SMOTESynthetic Minority Oversampling Technique
KNNK-Nearest Neighbors
IMFInternational Monetary Fund

References

  1. Reinhart, C.M.; Rogoff, K.S. Financial and Sovereign Debt Crises: Some Lessons Learned and Those Forgotten. J. Bank. Financ. Econ. 2015, 4, 5–17. Available online: https://www.ceeol.com/search/article-detail?id=696050 (accessed on 30 May 2025). [CrossRef]
  2. Laeven, L.; Valencia, F. Systemic Banking Crises Database II. IMF Econ. Rev. 2020, 68, 307–361. [Google Scholar] [CrossRef]
  3. Babecký, J.; Havránek, T.; Matějů, J.; Rusnák, M.; Šmídková, K.; Vašíček, B. Banking, Debt, and Currency Crises in Developed Countries: Stylized Facts and Early Warning Indicators. J. Financ. Stab. 2014, 15, 1–17. [Google Scholar] [CrossRef]
  4. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. Available online: https://www.jstor.org/stable/2699986 (accessed on 28 June 2026). [CrossRef]
  5. Laeven, L.; Valencia, F. Systemic Banking Crises Database. IMF Econ. Rev. 2013, 61, 225–270. [Google Scholar] [CrossRef]
  6. Reimann, C. Predicting financial crises: An evaluation of machine learning algorithms and model explainability for early warning systems. Rev. Evol. Political Econ. 2024, 5, 51–83. [Google Scholar] [CrossRef]
  7. Bluwstein, K.; Buckmann, M.; Joseph, A.; Kapadia, S.; Şimşek, Ö. Credit Growth, the Yield Curve and Financial Crisis Prediction: Evidence from a Machine Learning Approach. J. Int. Econ. 2023, 145, 103773. [Google Scholar] [CrossRef]
  8. Reinhart, C.M.; Rogoff, K.S.; Trebesch, C.; Reinhart, V.R. Global Crises Data by Country; Harvard Business School: Boston, MA, USA, 2019. [Google Scholar]
  9. Kaminsky, G.; Lizondo, S.; Reinhart, C.M. Leading indicators of currency crises. Staff. Pap. 1998, 45, 1–48. [Google Scholar] [CrossRef]
  10. Drehmann, M.; Juselius, M. Evaluating early warning indicators of banking crises: Satisfying policy requirements. Int. J. Forecast. 2014, 30, 759–780. [Google Scholar] [CrossRef]
  11. Betz, F.; Opricǎ, S.; Peltonen, T.A.; Sarlin, P. Predicting distress in European banks. J. Bank. Financ. 2014, 45, 225–241. [Google Scholar] [CrossRef]
  12. Laeven, L. Banking crises: A review. Annu. Rev. Financ. Econ. 2011, 3, 17–40. [Google Scholar] [CrossRef]
  13. Idris, I.T.; Nayan, S. The Joint Effects of Oil Price Volatility and Environmental Risks on Non-Performing Loans: Evidence from Panel Data of Organisation of the Petroleum Exporting Countries. Int. J. Energy Econ. Policy 2016, 6, 522–528. Available online: https://dergipark.org.tr/en/pub/ijeeep/article/351108?issue_id=31918 (accessed on 12 June 2026).
  14. Atoi, N.V. Non-Performing Loan and Its Effects on Banking Stability: Evidence from National and International Licensed Banks in Niger. CBN J. Appl. Stat. (JAS) 2018, 9, 43–74. [Google Scholar] [CrossRef]
  15. Adegbite, E. Good Corporate Governance in Nigeria: Antecedents, Propositions and Peculiarities. Int. Bus. Rev. 2015, 24, 319–330. [Google Scholar] [CrossRef]
  16. Alley, I.; Hassan, H.; Wali, A.; Suleiman, F. Banking Sector Reforms in Nigeria: An Empirical Appraisal. J. Financ. Regul. Compliance 2023, 31, 351–378. [Google Scholar] [CrossRef]
  17. Iwenya, R.; Adeghe, R. Macroeconomic Volatility and Bank Soundness in Nigeria. J. Glob. Account. 2026, 12, 66–82. Available online: https://journals.unizik.edu.ng/joga (accessed on 13 June 2026).
  18. Eke, P.O.; Achugamonu, B.U.; Yunisa, S.; Osuma, G.O. Macroeconomic Risks and Financial Sector Stability: The Nigerian Case. Decision 2020, 47, 233–249. [Google Scholar] [CrossRef]
  19. James, G.; Witten, D.M.; Hastie, T.; Tibshirani, R. An Introduction to Statistical Learning: With Applications in R; Springer Texts in Statistics; Springer: New York, NY, USA, 2013. [Google Scholar] [CrossRef]
  20. Beutel, J.; List, S.; von Schweinitz, G. Does machine learning help us predict banking crises? J. Financ. Stab. 2019, 45, 100693. [Google Scholar] [CrossRef]
  21. Garcia-Lopez, Y.J.; Marquez, P.H.; Morales, N.N. Microfinance Institutions Failure Prediction in Emerging Countries, a Machine Learning Approach. PLoS ONE 2025, 20, e0321989. [Google Scholar] [CrossRef] [PubMed]
  22. Gogas, P.; Papadimitriou, T.; Agrapetidou, A. Forecasting bank failures and stress testing: A machine learning approach. Int. J. Forecast. 2018, 34, 440–455. [Google Scholar] [CrossRef]
  23. Bussmann, N.; Giudici, P.; Marinelli, D.; Papenbrock, J. Explainable machine learning in credit risk management. Comput. Econ. 2021, 57, 203–216. [Google Scholar] [CrossRef]
  24. Liu, L.; Chen, C.; Wang, B. Predicting Financial Crises with Machine Learning Methods. J. Forecast. 2022, 41, 871–910. [Google Scholar] [CrossRef]
  25. Yin, S. Machine Learning Algorithms for Early Warning Systems: Predicting Systemic Financial Crises through Non-Linear Econometric Models. Int. J. Econ. Financ. Stud. 2024, 16, 286–311. [Google Scholar]
  26. Trickz, J. African Crises Dataset. Available online: https://www.kaggle.com/jaimetrickz/african-crises (accessed on 16 May 2025).
Figure 1. Correlation heatmap of macroeconomic indicators.
Figure 1. Correlation heatmap of macroeconomic indicators.
Analytics 05 00028 g001
Figure 2. Feature importance for detecting banking crises in Nigeria. Importance scores are computed using gain-based importance from the Gradient Boosting model. Features with importance exceeding 0.01 were selected for model training. The colour intensity of the bars reflects the magnitude of the importance score, with darker blue shades indicating higher importance.
Figure 2. Feature importance for detecting banking crises in Nigeria. Importance scores are computed using gain-based importance from the Gradient Boosting model. Features with importance exceeding 0.01 were selected for model training. The colour intensity of the bars reflects the magnitude of the importance score, with darker blue shades indicating higher importance.
Analytics 05 00028 g002
Figure 3. Precision–Recall curves for all models.
Figure 3. Precision–Recall curves for all models.
Analytics 05 00028 g003
Figure 4. ROC curves for all models.
Figure 4. ROC curves for all models.
Analytics 05 00028 g004
Figure 5. Crisis detection timeline: estimated crisis probabilities and observed banking crises. The vertical dashed lines indicate the years of the banking crisis (observed) while the horizontal dashed line indicates the alert threshold (0.5), above which the probabilities are considered as high risk (Red zone).
Figure 5. Crisis detection timeline: estimated crisis probabilities and observed banking crises. The vertical dashed lines indicate the years of the banking crisis (observed) while the horizontal dashed line indicates the alert threshold (0.5), above which the probabilities are considered as high risk (Red zone).
Analytics 05 00028 g005
Table 1. Overview of strategies for banking crisis detection.
Table 1. Overview of strategies for banking crisis detection.
Author/YearMethodology TypeKey FeaturesPerformance MetricsLimitations
[9]Traditional EWS: Logit/ProbitSignal extraction approachFalse alarm ratesLinear assumptions; high false positives
[1]Historical analysisCrisis incidence analysisFrequency analysisDescriptive only; non-operational
[2,12]Systemic crisis databaseCrisis datingSystemic risk indicatorsRetrospective; limited tracking horizon
[10]Credit-to-GDP gapSingle-indicator EWSSignal extractionCountry-specific limitations
[11]Multivariate logitStress testingAUROC, CalibrationData requirements for emerging markets
[4]XGBoostEnsemble learningAccuracyBlack-box nature
[22]SVM, Random ForestKernel methodsAUROCLimited emerging economy focus
[20]ML vs TraditionalComparative analysisAUROCData constraints
[21]Ensemble methodsRandom ForestF1-score, AccuracyInterpretability challenges
[6]XGBoost, Random ForestEnsemble learningAUROC, F1-scoreEuropean focus only
[23]XAI with SHAPExplainable AISHAP valuesAccuracy–interpretability trade-off
[3]Hybrid EWSBayesian methodsclassification accuracyImplementation complexity
Table 2. Feature selection summary: candidate variables, importance scores, and selection status.
Table 2. Feature selection summary: candidate variables, importance scores, and selection status.
FeatureImportance ScoreSelected
exch_usd0.425Yes
year0.305Yes
systemic_crisis0.139Yes
inflation_annual_cpi0.131Yes
case0.000No
sovereign_external_debt_default0.000No
domestic_debt_in_default0.000No
gdp_weighted_default0.000No
independence0.000No
currency_crises0.000No
inflation_crises0.000No
Note: Selection threshold: importance > 0.01.
Table 3. Key variable summary statistics.
Table 3. Key variable summary statistics.
VariableCountMeanStdMin25%50%75%Max
case6045.00.045.045.045.045.045.0
year601983.8817.881954.01968.751983.51999.252014.0
systemic_crisis600.1670.3760.00.00.00.01.0
exch_usd6038.9559.080.00.00.78100.85158.27
domestic_debt_in_default600.00.00.00.00.00.00.0
sovereign_external_debt_default600.150.360.00.00.00.01.0
gdp_weighted_default600.00.00.00.00.00.00.0
inflation_annual_cpi6014.7715.65−4.556.1010.7616.3072.73
independence600.90.300.01.01.01.01.0
currency_crises600.1670.3760.00.00.00.01.0
inflation_crises600.20.4030.00.00.00.01.0
banking_crisis600.1830.3900.00.00.00.01.0
Table 4. Cross-validation performance metrics on the SMOTE-resampled training set (AUC refers to ROC-AUC).
Table 4. Cross-validation performance metrics on the SMOTE-resampled training set (AUC refers to ROC-AUC).
ModelAccuracyF1-ScorePrecisionRecallROC-AUC
Logistic Regression0.78080.75690.87310.73930.9638
Random Forest0.92330.92700.92000.95000.9953
SVM0.57420.46820.49810.50000.8821
XGBoost0.87170.86870.89500.86790.9828
Note: Acc. = Accuracy; Prec. = Precision; Rec. = Recall.
Table 5. Evaluation metrics on the test set with bootstrapped 95% confidence intervals. Class 1 corresponds to banking crises (AUC refers to ROC-AUC).
Table 5. Evaluation metrics on the test set with bootstrapped 95% confidence intervals. Class 1 corresponds to banking crises (AUC refers to ROC-AUC).
ModelAcc.F1 (0)F1 (1)Prec. (1)Rec. (1)AUC
Logistic Regression0.830.890.670.501.001.00 [1.00, 1.00]
Random Forest0.750.820.570.401.001.00 [1.00, 1.00]
SVM0.830.890.670.501.001.00 [1.00, 1.00] 1
XGBoost0.920.950.800.671.001.00 [1.00, 1.00]
1 The SVM model initially yielded an ROC-AUC of 0.00 because the probability outputs were correlated with the reverse orientation of the classes at the time of testing; non-crisis observations were assigned higher estimated probabilities than crisis observations. This problem influenced the ranking applied in ROC-AUC calculation but had no significant effect on the classification accuracy or recall at the decided level. The ROC-AUC value was also corrected by correcting the class-label orientation in the evaluation pipeline. The sensitivity of probability calibration and performance measures in very small samples is also a concern in this context. Note: Acc. = Accuracy; Prec. = Precision; Rec. = Recall.
Table 6. Test set performance comparison: models with and without the “year” variable.
Table 6. Test set performance comparison: models with and without the “year” variable.
ModelYearAcc.F1 (1)Prec. (1)Rec. (1)AUC
Logistic RegressionWith0.830.670.501.001.00
Logistic RegressionWithout0.830.670.501.001.00
Random ForestWith0.750.570.401.001.00
Random ForestWithout0.830.670.501.001.00
SVMWith0.830.670.501.001.00
SVMWithout0.830.670.501.001.00
XGBoostWith0.920.800.671.001.00
XGBoostWithout0.920.800.671.001.00
Note: Acc. = Accuracy; Prec. = Precision; Rec. = Recall.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Mandiwana, N.; Ravele, T.; Sigauke, C.; Netshikweta, R. Machine Learning-Based Crisis Detection Framework for Banking Systems: A Case Study of Nigeria. Analytics 2026, 5, 28. https://doi.org/10.3390/analytics5030028

AMA Style

Mandiwana N, Ravele T, Sigauke C, Netshikweta R. Machine Learning-Based Crisis Detection Framework for Banking Systems: A Case Study of Nigeria. Analytics. 2026; 5(3):28. https://doi.org/10.3390/analytics5030028

Chicago/Turabian Style

Mandiwana, Ntanganedzeni, Thakhani Ravele, Caston Sigauke, and Rendani Netshikweta. 2026. "Machine Learning-Based Crisis Detection Framework for Banking Systems: A Case Study of Nigeria" Analytics 5, no. 3: 28. https://doi.org/10.3390/analytics5030028

APA Style

Mandiwana, N., Ravele, T., Sigauke, C., & Netshikweta, R. (2026). Machine Learning-Based Crisis Detection Framework for Banking Systems: A Case Study of Nigeria. Analytics, 5(3), 28. https://doi.org/10.3390/analytics5030028

Article Metrics

Back to TopTop