Next Article in Journal
Credit Risk Index as a Support Tool for the Financial Inclusion of Smallholder Coffee Producers
Next Article in Special Issue
Moderating Role of Risk Management Committee on Board of Directors’ Characteristics and Corporate Risk Disclosure Nexus: Emerging Market Evidence
Previous Article in Journal
Aligning Inclusive Finance with the European Union’s Digital–Green Twin Transition
Previous Article in Special Issue
Regulating Green Finance and Managing Environmental Risks in the Conditions of Global Uncertainty
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Predicting Financial Contagion: A Deep Learning-Enhanced Actuarial Model for Systemic Risk Assessment

by
Khalid Jeaab
1,*,
Youness Saoudi
2,
Smaaine Ouaharahe
3 and
Moulay El Mehdi Falloul
1
1
Economics and Management Laboratory, Sultan Moulay Slimane University, Khouribga 25000, Morocco
2
Advanced Systems Engineering Laboratory, Ibn Tofail University, Kenitra 14000, Morocco
3
Organization Economics and Management Laboratory, Ibn Tofail University, Kenitra 14000, Morocco
*
Author to whom correspondence should be addressed.
J. Risk Financ. Manag. 2026, 19(1), 72; https://doi.org/10.3390/jrfm19010072
Submission received: 25 June 2025 / Revised: 1 August 2025 / Accepted: 11 August 2025 / Published: 16 January 2026
(This article belongs to the Special Issue Financial Regulation and Risk Management amid Global Uncertainty)

Abstract

Financial crises increasingly exhibit complex, interconnected patterns that traditional risk models fail to capture. The 2008 global financial crisis, 2020 pandemic shock, and recent banking sector stress events demonstrate how systemic risks propagate through multiple channels simultaneously—e.g., network contagion, extreme co-movements, and information cascades—creating a multidimensional phenomenon that exceeds the capabilities of conventional actuarial or econometric approaches alone. This paper addresses the fundamental challenge of modeling this multidimensional systemic risk phenomenon by proposing a mathematically formalized three-tier integration framework that achieves 19.2% accuracy improvement over traditional models through the following: (1) dynamic network-copula coupling that captures 35% more tail dependencies than static approaches, (2) semantic-temporal alignment of textual signals with network evolution, and (3) economically optimized threshold calibration reducing false positives by 35% while maintaining 85% crisis detection sensitivity. Empirical validation on historical data (2000–2023) demonstrates significant improvements over traditional models: 19.2% increase in predictive accuracy (R2 from 0.68 to 0.87), 2.7 months earlier crisis detection compared to Basel III credit-to-GDP indicators, and 35% reduction in false positive rates while maintaining 85% crisis detection sensitivity. Case studies of the 2008 crisis and 2020 market turbulence illustrate the model’s ability to identify subtle precursor signals through integrated analysis of network structure evolution and semantic changes in regulatory communications. These advances provide financial regulators and institutions with enhanced tools for macroprudential supervision and countercyclical capital buffer calibration, strengthening financial system resilience against multifaceted systemic risks.

1. Introduction

The stability of the global financial system is currently facing unprecedented challenges. The increasing complexity of financial markets, the acceleration of information flows, the growing interconnectedness of institutions, and the emergence of new risks are profoundly transforming the global financial landscape. Successive crises—from the collapse of 2008 to the turmoil of the 2020 pandemic—have clearly demonstrated the limitations of traditional approaches to financial risk monitoring and modeling.
These events have revealed a worrying reality: our conventional tools for measuring and anticipating systemic risks remain inadequate in the face of the complexity and non-linearity of contemporary financial dynamics. The economic and social cost of this inadequacy is considerable. According to the IMF, the 2008 crisis led to a cumulative loss of global output exceeding USD 10 trillion. Beyond their immediate financial impact, these crises erode confidence in institutions and exacerbate socioeconomic inequalities.
In this context, the convergence between actuarial science—with its mathematical rigor and tradition of prudent risk modeling—and recent advances in Big Data offers a major opportunity to fundamentally rethink our understanding and management of systemic risks.
On the one hand, actuarial science has developed a robust conceptual framework for quantifying uncertainty over the centuries. Based on rigorous statistical principles and well-defined parametric assumptions, it excels at modeling phenomena whose stochastic properties are relatively stable and well understood. Its deductive approach, using theoretical models to interpret data, promotes interpretability and mathematical consistency.
On the other hand, Big Data and machine learning technologies offer a radically different paradigm. Their inductive, data-driven approach, capable of identifying complex structures without a priori assumptions, offers exceptional flexibility in dynamic environments and non-parametric relationships. These methods can exploit massive volumes of heterogeneous and unstructured data that are inaccessible to conventional models.
However, this duality also reveals fundamental tensions. Actuarial rigor becomes restrictive in the face of highly non-linear emerging phenomena characterized by complex interactions. Conversely, Big Data approaches raise legitimate questions about their interpretability, stability, and causal relevance—essential dimensions in a regulatory and prudential context. Integrating these two epistemological worlds is not simply a technical challenge, but a strategic necessity for the future of financial risk modeling. The goal is not to replace one paradigm with the other, but to build a harmonious synthesis that capitalizes on their complementary strengths.
Our research revolves around a central question: how can we design a methodological framework that coherently and operationally integrates the fundamentals of actuarial science with the advanced capabilities of Big Data to significantly improve the detection, quantification, and management of systemic risks?
This general question can be broken down into several specific questions that structure our approach:
  • How can we reconcile deductive and inductive approaches in a unified framework that preserves both the theoretical rigor and empirical flexibility necessary to understand systemic risks?
  • What methodological architectures can integrate traditional financial data with alternative sources (textual, transactional, and geospatial) to capture the multidimensionality of systemic risks?
  • How can we maintain an optimal balance between predictive performance and interpretability in a context where regulatory decisions require transparency and justification?
  • What empirical validation mechanisms are appropriate for evaluating models designed to anticipate rare and heterogeneous events such as systemic crises?
Research Questions Solutions Framework:
Our methodological framework directly addresses each research question through specific innovations:
Q1—Deductive/Inductive reconciliation: We present the construction of our composite indicator using a gradient boosting framework that preserves actuarial prudence principles while leveraging the pattern detection capabilities of machine learning.
Q2—Multi-source data integration: We detail our approach to integrating text analysis, combining TF-IDF transformation with network centrality measures to capture multidimensional risk signals.
Q3—Performance–interpretability balance: We introduce our Alert Threshold Optimization methodology using an economic cost function with F-beta optimization to maintain transparency while maximizing predictive performance.
Q4—Rare events validation: We present our Model Performance Evaluation using temporal cross-validation with bootstrap iterations, maintaining normal/crisis ratios of 88:12 to ensure robust validation despite limited crisis observations.
To answer these questions, we propose an innovative three-level actuarial framework that combines data preparation, hybrid modeling, and validation/interpretation of the results. Our approach differs from previous work in several key ways:
First, unlike approaches that simply juxtapose different methodologies, we propose an integrative framework that mathematically formalizes the interactions between actuarial models and Big Data techniques, enabling cross-fertilization rather than mere coexistence.
Second, we specifically develop an in-depth application to systemic risk modeling, an area where the integration of actuarial and Big Data approaches remains in its infancy despite its crucial importance for financial stability.
Third, our framework pays particular attention to the interpretability and actionability of results in a regulatory context, a dimension often overlooked in work focused primarily on pure predictive performance.
Our specific objectives include the following:
  • Developing a modular architecture that allows for the gradual integration of Big Data methods into existing actuarial models;
  • Mathematically formalizing this approach to ensure its theoretical consistency and generality;
  • Empirically demonstrating its predictive superiority over historical episodes of financial stress;
  • Proposing concrete applications for regulators and financial institutions.
This research is organized according to a logical progression that covers the evolution of systemic risk concepts, traditional modeling approaches, and recent advances in Big Data. The proposed conceptual and methodological framework is illustrated by a detailed case study on systemic risk modeling, including a historical analysis of the 2008 and 2020 crises.
We validate our approach using concrete case studies covering different manifestations of systemic risk, thereby demonstrating its generality and robustness. The implications for macroprudential regulation and financial stability are analyzed, as well as the associated technical, ethical, and regulatory challenges.
Through this approach, we hope to make a significant contribution at both the theoretical and operational levels to improving the forecasting and management of systemic risks, thereby strengthening the resilience of the financial system in the face of future challenges in this rapidly evolving field.

2. Literature Review

2.1. Systemic Risk Evolution

The concept of systemic risk, although omnipresent in contemporary discussions on financial stability, has undergone significant evolution in its conceptualization and formalization. Historically, the term mainly referred to the risk of a chain reaction of failures among financial institutions, but its definition has broadened and become more precise in the wake of successive financial crises.
According to the now widely accepted definition proposed by the Financial Stability Board (FSB, 2010), systemic risk represents “the risk of disruption to financial services caused by a deterioration in all or part of the financial system, with the potential to generate serious adverse consequences for the real economy.” This definition highlights two key dimensions of systemic risk: its endogenous nature within the financial system and its potential impact on the real economy.
De Bandt and Hartmann (2000) made a fundamental distinction between “strong” and “weak” systemic shocks, with the former leading to the failure of solvent institutions through contagion, while the latter causes a widespread transmission of shocks without necessarily leading to cascading failures. This conceptualization has been enriched by the work of Allen and Gale (2000), who formalized the mechanisms of direct contagion (bilateral exposures) and indirect contagion (asset devaluation and liquidity spirals).
Recent developments in the literature, notably the contributions of Brunnermeier et al. (2009) and Adrian and Brunnermeier (2016), have highlighted the distinctive characteristics of systemic risk compared to traditional financial risks: non-linearity with threshold effects and feedback loops that amplify initial shocks, endogeneity (where risk emerges from interactions between agents and is not simply imposed from outside), the complexity of interconnections that create vulnerabilities that are not apparent at the individual level, procyclicality (where amplification mechanisms tend to reinforce financial cycles), and multidimensionality, combining credit, market, liquidity, and operational risks.

2.2. Traditional vs. Big Data Approaches

The conceptualization of systemic risk has undergone major changes in response to financial crises. Kindleberger (1978) and Minsky (1992) had already proposed theoretical models describing the phases of expansion, euphoria, distress, and panic characteristic of financial crises. However, these approaches remained primarily descriptive and qualitative.
The Asian crisis of 1997–1998 highlighted the importance of “balance sheet effects” and exchange rate asymmetries as factors of systemic risk Goldstein (1998), while the collapse of the LTCM fund in 1998 revealed the dangers posed by excessive leverage and concentrated positions President’s Working Group on Financial Markets (1999).
The global financial crisis of 2007–2009 was a major turning point, highlighting new dimensions of systemic risk that had previously been underestimated: the role of shadow banking in the creation and propagation of risk Gorton and Metrick (2012), the importance of interconnections in short-term financing markets Brunnermeier (2009), the risks associated with the complexity and opacity of structured finance products Coval et al. (2009), and the impact of misaligned incentives and moral hazard on excessive risk-taking Bebchuk and Spamann (2010).
More recently, the financial turmoil associated with the COVID-19 pandemic in 2020 highlighted the importance of exogenous non-financial shocks and pre-existing vulnerabilities in the financial system Darracq Parièset al. (2021), while the March 2023 episode involving Silicon Valley Bank highlighted the risks associated with sector concentration and maturity transformation Kashyap et al. (2023).
Early attempts to quantify systemic risk relied mainly on aggregate macroeconomic and financial indicators. These approaches, summarized by Borio and Lowe (2002) and Borio and Drehmann (2009), focused on identifying macrofinancial imbalances as precursors to crises: deviations in credit from long-term trends, real estate price/income ratios, indicators of external imbalances (current account deficits), and measures of monetary expansion.
These indicators have the advantage of simplicity and a certain transparency but suffer from significant limitations, notably their inability to capture the complexity of interactions between financial institutions and markets Galati and Moessner (2013). Their calibration also remains problematic, with signals often ambiguous or delayed.
From a more theoretical perspective, dynamic stochastic general equilibrium (DSGE) models have been enhanced to incorporate the financial frictions and amplification mechanisms that characterize systemic crises. The seminal work of Bernanke et al. (1999) on the financial accelerator formalized how balance sheet constraints amplify initial macroeconomic shocks. This approach was extended by Gertler and Kiyotaki (2010) to explicitly model frictions in interbank markets and by Christiano et al. (2014) to incorporate endogenous default risk. Although these models provide a coherent theoretical framework, they often rely on simplifying assumptions about agent homogeneity and perfect rationality, and struggle to capture the non-linearities and extreme behaviors that characterize systemic crises (Stiglitz, 2018).
A significant body of literature has focused on explicitly modeling financial contagion mechanisms. The work of Allen and Gale (2000) established a theoretical framework distinguishing different network structures (complete, incomplete, and ring) and their respective resilience to liquidity shocks. This framework was expanded upon by Freixas et al. (2000) to incorporate uncertainty about the solvency of counterparties and then by Gai and Kapadia (2010), who modeled default cascades in complex financial networks.

2.3. Theoretical Foundations

The transition to Big Data-based approaches was initially marked by the development of systemic measures derived from high-frequency market data. These measures aim to quantify each institution’s contribution to overall systemic risk.
The CoVaR (Conditional Value-at-Risk) proposed by Adrian and Brunnermeier (2016) measures the value-at-risk of the financial system conditional on the distress of a specific institution. This approach was complemented by the MES (Marginal Expected Shortfall) of Acharya et al. (2017), which assesses an institution’s marginal contribution to the expected shortfall of the system. Brownlees and Engle (2017) developed SRISK, which quantifies an institution’s expected capital shortfall in the event of a systemic crisis, while Diebold and Yılmaz (2014) proposed an approach based on variance decomposition to measure spillovers between financial institutions.
The application of complex network theory to financial analysis represents a major methodological advance, enabling more accurate modeling of interconnections between institutions. The pioneering work of Battiston et al. (2012) introduced the concept of “DebtRank,” a recursive measure of systemic importance inspired by Google’s PageRank algorithm.
The application of machine learning techniques to the early detection of systemic crises is a recent and promising development. Unlike traditional early warning systems based on predefined thresholds, these approaches can capture complex non-linear interactions between explanatory variables.
Alessi and Detken (2018) demonstrated the superiority of random forests over traditional logistic models for predicting systemic banking crises, with a significant improvement in the area under the ROC curve. In a similar vein, Beutel et al. (2019) used ensemble methods (boosting) to identify the most relevant variables for detecting systemic vulnerabilities.
The exploitation of massive textual data represents a particularly promising frontier for the detection of emerging systemic risks. The wealth of information contained in communications from financial authorities and market participants offers valuable signals that complement traditional quantitative indicators.
Actuarial science has historically developed around fundamental principles that make it a particularly suitable framework for financial risk management. Its deductive approach, based on well-established theoretical models, emphasizes mathematical consistency and interpretability. Actuarial methods excel at quantifying uncertainty through well-defined probability distributions, assessing risks in long-term contexts, and integrating regulatory and prudential constraints.

2.4. Research Gap Identification

This Table 1 summarizes the methodological innovations introduced by this study compared to existing approaches in four key areas. The quantified improvements are particularly impressive, notably the 19.2% gain in accuracy thanks to the mathematical formalization of actuarial-ML interactions, increasing from an R 2 of 0.68 to 0.87. The most remarkable innovation concerns network analysis with dynamic network-copula coupling, which improves the capture of tail dependencies by 35%, significantly exceeding traditional static assumptions. The temporal integration of text with network evolution enables early detection of 2.7 months, while economic optimization of thresholds reduces false positives by 35%, demonstrating a holistic approach that combines mathematical rigor and practical relevance.
The principle of prudence, central to the actuarial approach, requires safety margins in risk assessment and favors conservative assumptions. This philosophy is particularly relevant in the context of systemic risk, where the consequences of underestimation can be catastrophic for the entire financial system.
Traditional actuarial models are based on strong parametric assumptions concerning probability distributions, risk independence, and parameter stability over time. While these assumptions facilitate interpretation and regulatory control, they can prove restrictive in the face of the complexity and non-linearity of contemporary systemic phenomena.
Big Data approaches take a radically different philosophy, favoring induction and the discovery of patterns in data rather than the validation of pre-established theoretical assumptions. This approach is characterized by its ability to process massive volumes of heterogeneous data, identify complex relationships without a priori assumptions, adapt dynamically to structural changes, and exploit non-traditional sources of information.
Machine learning techniques, at the heart of the Big Data paradigm, excel at detecting weak signals and emerging patterns, managing high dimensionality, adapting to non-stationary environments, and integrating multimodal data (numerical, textual, and temporal).
However, this flexibility comes with significant challenges in terms of the interpretability of “black box” models, the stability of predictions in the face of data variations, the risk of overfitting on specific patterns, and difficulties in establishing causal relationships.
The traditional actuarial approach is based on several theoretical pillars that are particularly relevant to systemic risk analysis. Extreme value theory, through the Pickands-Balkema-de Haan framework, provides a solid basis for modeling the rare and catastrophic events characteristic of systemic crises Embrechts et al. (1997). Copula theory, developed by Sklar (1959) and extended by Joe (1997) and Nelsen (2006), allows for flexible modeling of multivariate dependence structures, which is particularly crucial for capturing extreme co-movements during periods of crisis.
Bayesian credibility integrates prior knowledge with empirical data in a Bayesian framework, improving statistical inference in data-limited contexts typical of systemic events Bühlmann and Gisler (2005). Collective ruin theory, through Lundberg-Cramér models, provides a framework for analyzing simultaneous or cascading failures of financial institutions Asmussen and Albrecher (2010).
Recent advances in data science complement the actuarial approach with unsupervised learning models for anomaly detection and clustering, enabling the identification of emerging patterns of systemic vulnerability without prior assumptions about the form of these patterns Chandola et al. (2009).
Natural language processing, with word embeddings and attention models Vaswani et al. (2017), can extract early stress signals from unstructured textual data such as financial communications. Recurrent neural networks and long short-term memory networks capture the complex temporal dependencies and long memory structures characteristic of financial time series Hochreiter and Schmidhuber (1997). Ensemble models and boosting provide increased robustness and superior predictive power by combining multiple models (Friedman, 2001).
The convergence of the actuarial and Big Data frameworks creates significant theoretical synergies. Epistemological complementarity enables the actuarial approach, based on parametric models with a strong theoretical foundation, to complement the more flexible and adaptive, but sometimes less interpretable, Big Data approach. The integration of the two approaches mitigates the biases inherent in each, notably the specification bias of parametric models and the overlearning bias of machine learning techniques.
This hybrid approach simultaneously models the idiosyncratic and systemic components of risk, reflecting their interdependence in the real world. Financial network analysis uses graph theory to model interconnections between institutions, where topology determines resilience to shocks. Measures of centrality provide quantitative indicators of the systemic importance of institutions (Battiston et al., 2012), complementing size-based regulatory approaches.

2.4.1. Modeling Extreme Dependencies and Textual Analysis

Dependencies between financial institutions typically intensify in times of crisis, requiring specific tools. Tail dependency coefficients, introduced by Joe (1997), quantify the conditional probability of observing simultaneous extreme values, providing a measure of potential contagion in times of crisis.
The copula theoretical framework offers a flexible modeling of dependence structures independent of marginal distributions. Sklar’s theorem states that a multivariate distribution can be decomposed into marginal distributions and a copula function capturing the dependence structure Sklar (1959).
Financial communication analysis uses Shannon’s information theory to formalize how uncertainty and entropy in communications can reveal information about the underlying state of the system. Pre-trained language models and transfer learning make it possible to exploit general linguistic knowledge while capturing the specificities of financial discourse.
The financial system can be understood as a complex adaptive system with emergent properties that cannot be deduced from an isolated analysis of its components Farmer et al. (2012). Financial agents adapt their behavior according to the environment and the actions of other agents, creating evolutionary dynamics Arthur (2014). Positive feedback mechanisms can amplify small disturbances, creating disproportionate effects (Sornette, 2003).
This complex nature imposes methodological constraints: traditional approaches assuming equilibrium and complete rationality do not adequately capture the out-of-equilibrium dynamics characteristic of crises (Bookstaber, 2017). Effective modeling of systemic risk requires the integration of processes operating at different temporal and organizational scales (Battiston et al., 2016).
The first major limitation concerns the persistent segmentation between different methodological approaches. Models based on financial networks, measures based on market data, and textual analyses are generally developed in parallel, with little effort to integrate them into a unified framework (Bisias et al., 2012). This fragmentation limits the ability to capture the complex interactions between different dimensions of systemic risk.
Current models focus primarily on financial transmission and contagion channels, often neglecting interactions with the real economy, behavioral dimensions, and institutional factors Cerutti et al. (2022). Integrating these non-financial dimensions represents a significant methodological challenge, but one that is crucial for a holistic understanding of systemic risk.

2.4.2. Existing Hybrid Model Limitations

Single-dimension integration approaches:
  • Beutel et al. (2019): Random forests for crisis prediction—combines variables but lacks a theoretical foundation;
  • Danielsson et al. (2022): Network analysis with ML—focuses only on topological features and ignores extreme dependencies;
  • Skokov and Chiriac (2022): Text analysis with traditional indicators—separate processing without mathematical integration.
Mathematical formalization gap:
Current approaches use ad-hoc combinations rather than theoretically grounded integration:
  • Ensemble methods (Alessi & Detken, 2018): Simple weighted averages without considering variable interactions:
    -
    Formula: R i s k e n s e m b l e = w i · I n d i c a t o r i (linear combination);
    -
    Limitation: No cross-fertilization between methodologies.
  • Parallel processing (Benoit et al., 2017): Separate models for different risk dimensions.
    -
    Network analysis→Risk Score1;
    -
    Market indicators→Risk Score2;
    -
    Text analysis→Risk Score3;
    -
    Limitation: No interaction terms or dependency modeling.
Our innovation—Mathematical integration:
  • Network–copula coupling: Λ i j = f ( A i j , S I I ( i ) , S I I ( j ) ) —tail dependencies informed by network structure;
  • Semantic-temporal alignment: Text sentiment weights adjusted by network centrality evolution;
  • Unified optimization: Single objective function incorporating all dimensions with economic cost weighting.
Quantified difference:
  • Existing: Independent processing→Simple aggregation;
  • This study: Mathematical cross-fertilization→19.2% performance improvement.
The growing use of sophisticated machine learning algorithms raises legitimate concerns about their interpretability and acceptability in a macroprudential policy context Danielsson and Shin (2003). Regulators generally favor models with explicit and understandable causal mechanisms, which can conflict with the inherent complexity of Big Data approaches.
Finally, rigorous empirical validation of systemic risk models is hampered by the rarity of observable systemic crises and their heterogeneity Reinhart and Rogoff (2009). This fundamental limitation complicates the comparative evaluation of different approaches and can lead to problems of overfitting on specific historical episodes.
Our research aims to fill these gaps by proposing an integrated framework that combines the theoretical foundations of actuarial science with the advanced analytical capabilities of Big Data. This hybrid approach makes it possible to simultaneously exploit structured and unstructured data, capture the complex interactions between financial institutions, and maintain a level of interpretability compatible with regulatory requirements.

3. Systematic Comparison of Systemic Risk Modeling Approaches

The proliferation of systemic risk modeling approaches over the past two decades necessitates a systematic evaluation of their relative strengths, limitations, and complementary potentials. This section provides comprehensive comparisons across three major methodological paradigms that inform our integrated framework design. The analysis draws from extensive literature reviews (Benoit et al., 2017; Bisias et al., 2012) and recent comparative studies (Alessi & Detken, 2018; Beutel et al., 2019) to establish empirical benchmarks for methodological selection.
Our comparative framework evaluates approaches across five key dimensions: (1) theoretical foundation and interpretability, (2) predictive accuracy and early warning capability, (3) computational requirements and scalability, (4) regulatory acceptability and transparency, and (5) robustness across different crisis types. Table 2, Table 3 and Table 4 synthesize these evaluations, providing the empirical foundation for our hybrid approach design.

3.1. Traditional Systemic Risk Modeling Approaches

Traditional approaches to systemic risk modeling have evolved through successive financial crises, each revealing new dimensions of systemic vulnerability while highlighting limitations of existing frameworks. The global financial crisis of 2007–2009 marked a watershed moment, exposing fundamental inadequacies in pre-crisis risk models and spurring methodological innovations across academic and policy institutions Brunnermeier (2009); Gorton and Metrick (2012).
Methodological assessment and limitations:
Aggregate macroeconomic indicators, while forming the backbone of many early warning systems, including the Basel III framework, suffer from fundamental timing issues. Drehmann et al. (2011) demonstrated that credit-to-GDP gaps, the Basel Committee’s preferred indicator, provide average lead times of only 12–15 months with false positive rates exceeding 40% in developed economies. The 2020 pandemic crisis further exposed these limitations, as traditional indicators failed to capture the rapid shift from real economy shock to financial system stress Darracq Parièset al. (2021).
Dynamic stochastic general equilibrium (DSGE) models, despite their theoretical elegance, face criticism for unrealistic behavioral assumptions during crisis periods. Stiglitz (2018) argues that equilibrium-based frameworks cannot capture the ”animal spirits” and coordination failures that characterize systemic crises. Empirical evidence supports this critique: DSGE-based crisis predictions show accuracy rates below 35% for major crisis episodes Christiano et al. (2014).
Financial contagion models represent a significant advance in capturing institutional interconnections but rely heavily on static network representations. Recent research by Battiston et al. (2016) demonstrates that network topology changes dramatically during crisis periods, with correlation structures shifting from 0.3–0.4 in normal times and from 0.7–0.8 during stress episodes. This dynamic evolution undermines models based on fixed network structures.
Market-based measures like CoVaR and SRISK offer real-time assessment capabilities but exhibit strong pro-cyclical tendencies. Adrian and Brunnermeier (2016) acknowledge that these measures tend to be low precisely when risks are building up, and high when crises have already materialized. The COVID-19 crisis exemplified this limitation, with market-based indicators spiking only after widespread disruptions became apparent (Danielsson et al., 2022).

3.2. Big Data and Machine Learning Approaches

The integration of machine learning techniques into systemic risk modeling represents a paradigmatic shift from theory-driven to data-driven approaches. This evolution reflects both the availability of massive financial datasets and the limitations of traditional econometric methods in capturing complex, non-linear risk relationships (Danielsson et al., 2022; L. Chen et al. (2023)).
Performance analysis and trade-offs:
The machine learning landscape for systemic risk modeling reveals a fundamental trade-off between predictive accuracy and interpretability. Neural networks consistently achieve the highest predictive performance, with studies reporting AUC scores exceeding 0.90 for crisis prediction tasks Liu and Zhang (2024). However, their black-box nature renders them unsuitable for regulatory applications requiring transparent decision-making processes.
Random forests provide an attractive middle ground, combining strong predictive performance (AUC typically 0.82–0.89) with moderate interpretability through feature importance measures and SHAP (SHapley Additive exPlanations) values. Beutel et al. (2019) demonstrated that random forest models can identify crisis precursors 18–24 months in advance, significantly outperforming traditional logistic regression approaches.
Gradient boosting methods, particularly XGBoost and LightGBM implementations, show exceptional performance in financial applications while maintaining reasonable interpretability through tree visualization techniques. T. Chen and Guestrin (2016) report that gradient boosting achieves optimal bias–variance trade-offs for financial time series, explaining its adoption by major financial institutions for risk modeling applications.
Network analysis represents a unique category, sacrificing some predictive accuracy for exceptional interpretability and regulatory acceptance. Centrality measures like DebtRank (Battiston et al., 2012) and eigenvector centrality provide intuitive metrics for systemic importance that directly inform regulatory capital requirements and supervision priorities.
Regulatory acceptability assessment:
Regulatory suitability varies dramatically across machine learning approaches. The European Central Bank’s Macroprudential Bulletin Lo Duca et al. (2017) emphasizes that models used for policy decisions must provide clear economic interpretation and robust theoretical foundation.” This requirement effectively excludes neural networks and support vector machines from macroprudential applications.
The Federal Reserve’s Supervisory Guidance on Model Risk Management (SR 11-7) requires that models used for supervisory purposes demonstrate “conceptual soundness” and “appropriate ongoing monitoring.” These criteria favor interpretable approaches like network analysis and gradient boosting over black-box methods.
Our selection of gradient boosting and network analysis reflects this regulatory landscape while maximizing predictive performance within interpretability constraints. The combination achieves AUC scores of 0.87 (comparable to neural networks) while maintaining full transparency for regulatory review and validation.

3.3. Text Analysis Techniques for Financial Communications

The exploitation of textual data in financial risk assessment represents one of the most promising frontiers in systemic risk modeling. Financial communications contain forward-looking information not captured in traditional quantitative indicators, potentially providing crucial early warning signals Baker et al. (2016); Skokov and Chiriac (2022).
Methodological evolution and comparative performance:
Dictionary-based approaches, pioneered by Loughran and McDonald (2011), established the foundation for financial text analysis by recognizing that general sentiment lexicons poorly capture financial communication nuances. Words like “liability,” “volatile,” and “exposure” carry specific meanings in financial contexts that differ from general usage. The Loughran–McDonald financial lexicon achieves 65–75% accuracy in sentiment classification, with the crucial advantage of high temporal consistency (correlation stability r = 0.89 over rolling 12-month windows).
Supervised machine learning approaches, particularly support vector machines and random forests trained on labeled financial documents, achieve higher accuracy (8–85%) but require substantial training data and periodic retraining. Correa et al. (2021) demonstrate that machine learning models trained on Federal Reserve communications can predict policy changes with 82% accuracy, but performance degrades significantly when applied to communications from other central banks without retraining.
BERT and transformer models represent the current state-of-the-art in natural language processing, achieving 85–90% accuracy in financial sentiment classification. Azqueta-Gavaldón et al. (2023) show that fine-tuned BERT models can capture subtle semantic nuances in European Central Bank communications that escape traditional approaches. However, these models require substantial computational resources (GPU clusters) and show concerning temporal instability (correlation stability r = 0.68), limiting their practical applicability to real-time monitoring systems.
Hybrid approach design and validation:
Our hybrid multi-source approach combines the strengths of different techniques while mitigating individual weaknesses. The methodology incorporates the following:
  • Primary dictionary-based scoring: Using enhanced Loughran–McDonald lexicons with financial crisis-specific terms;
  • Machine learning validation: Cross-validation using random forest models trained on regulatory communications;
  • Expert panel review: Human validation of ambiguous classifications by financial economists;
  • Market-based anchoring: Correlation validation with observable market stress indicators (VIX and credit spreads).
This combination achieves 88–92% accuracy while maintaining high temporal consistency (r = 0.83), representing optimal performance for regulatory applications requiring both accuracy and stability.
Data source selection and bias mitigation:
Source selection critically impacts text analysis quality and potential biases. Our framework prioritizes official regulatory communications (60% weight) over media sources (40% weight) to minimize sensationalism bias. Primary sources include the following:
  • Central Bank communications: FOMC minutes, ECB press releases, and Bank of England Financial Stability Reports;
  • Regulatory announcements: Basel Committee publications and national supervisory statements;
  • Rating agency reports: Moody’s, S&P, Fitch sovereign, and banking assessments.
Secondary sources undergo bias detection algorithms identifying potentially misleading language patterns. Cross-source triangulation requires a minimum of three independent sources for sentiment validation, with expert panel review for conflicting assessments.

3.4. Synthesis and Methodological Integration Framework

The comparative analysis reveals complementary strengths across methodological paradigms that inform our integrated approach design. Traditional approaches provide a theoretical foundation and regulatory acceptance but lack predictive power and timeliness. Machine learning methods offer superior accuracy but sacrifice interpretability and stability. Text analysis captures forward-looking information but requires careful bias mitigation and source validation.
Integration principles:
Our hybrid framework addresses these limitations through systematic integration based on four principles:
  • Complementary information synthesis: Network measures capture structural vulnerabilities, extreme dependency models quantify tail risks, and text analysis provides early warning signals;
  • Interpretability preservation: Maintaining transparency through actuarial foundations while leveraging machine learning predictive power;
  • Robustness through diversification: Combining multiple methodologies reduces model risk and improves stability across different crisis types;
  • Regulatory compliance: Ensuring all components meet interpretability and validation requirements for macroprudential applications.
Optimal weight calibration:
Empirical validation determines optimal combination weights through cross-validation across crisis episodes:
  • Network analysis: 40% (structural vulnerability indicators);
  • Extreme dependencies: 35% (tail risk quantification);
  • Text analysis: 25% (forward-looking sentiment signals).
These weights maximize early detection capability (2.7 months advance warning) while maintaining high precision (88% accuracy) and regulatory acceptability (full interpretability through component analysis).
The integration framework represents a methodological advance beyond simple ensemble approaches, providing mathematically formalized interactions between actuarial theory and data-driven techniques. This synthesis enables cross-fertilization rather than mere coexistence, addressing the fundamental challenge of combining deductive and inductive paradigms in financial risk modeling.

4. Methodology and Model Construction

4.1. Network Analysis Framework

The financial network can be formalized as a weighted directed graph G = ( V , E , W ) where the following applies:
  • V = { 1 , 2 , , n } represents all n financial institutions;
  • E V , V represents all interbank exposures;
  • W : E R + is a function assigning a weight to each edge, representing exposure as a percentage of capital.
The weighted adjacency matrix A R n × n is defined as
A i j = w i j if ( i , j ) E 0 else
where w i j represents the exposure of institution i to institution j.
Representing the financial system as a graph captures the complex structure of relationships between institutions. In this formalism, the following applies:
  • Each node (financial institution) is an actor that can both influence others and be influenced by them;
  • Directed edges (exposures) represent potential channels of contagion;
  • Edge weights quantify the intensity of these channels.
The composite systemic importance index combines four complementary centrality measures:
  • Eigenvector centrality c E ( i ) : Captures recursive importance (“being connected to important nodes”);
  • Betweenness centrality c B ( i ) : Identifies bridge institutions between network communities;
  • In-degree centrality c i n ( i ) : Measures vulnerability to counterparty defaults;
  • Out-degree centrality c o u t ( i ) : Quantifies contagion propagation capacity.
The weighted combination is defined as
S I I ( i ) = α c E ( i ) + β c B ( i ) + γ c i n ( i ) + δ c o u t ( i )
where weights α   , β ,   γ ,   a n d   δ are empirically calibrated to maximize correlation with historical systemic impact measures, subject to the normalization constraint α + β + γ + δ = 1 .
Complete technical details are provided in Appendix A.

4.2. Extreme Dependency Modeling

Tail dependencies measure the tendency of institutions to suffer extreme losses simultaneously—a crucial phenomenon for understanding contagion in times of crisis.
Extreme dependencies between financial institutions intensify during crisis periods, requiring specialized modeling approaches. We employed a Student’s t–copula framework, which captures tail dependence unlike Gaussian alternatives.
The lower tail dependence coefficient λ L measures the conditional probability of simultaneous extreme losses between institutions. For n institutions, we construct a tail dependency matrix Λ where off-diagonal elements Λ i j = λ L ( X i , X j ) quantify pairwise extreme co-movement probabilities.
The Student’s t–copula with ν degrees of freedom and correlation matrix R is defined by
C ν , R ( μ 1 , μ 2 , , μ n ) = t ν , R ( t 1 ν ( μ 1 ) , t 1 ν ( μ 2 ) , , t 1 ν ( μ n ) )
Parameter estimation follows the two-step Inference Functions for Margins (IFMs) approach, with marginal parameters estimated first, followed by copula parameters. Complete technical details are provided in Appendix B.

4.2.1. Theoretical Foundation: Decision Theory and Signal Detection

The threshold optimization problem can be formalized within the Neyman–Pearson framework for hypothesis testing, where we test the null hypothesis H 0 (no crisis) against the alternative H 1 (crisis occurring). Unlike classical hypothesis testing with fixed significance levels, early warning applications require optimal threshold selection based on the economic consequences of decision errors.
Following Green and Swets (1966), we model the early warning system as a signal detection problem, where the following applies:
  • Signal: True crisis indicators in a noisy environment;
  • Noise: Normal market fluctuations and false indicators;
  • Decision rule: Threshold-based classification of system state;
  • Performance criterion: Expected economic cost minimization.
The optimal threshold emerges from the intersection of signal and noise distributions, weighted by relative costs and prior probabilities of crisis occurrence. This framework extends classical ROC analysis by incorporating economic consequences rather than purely statistical criteria.

4.2.2. Economic Cost Function Development

Optimal alert thresholds balance the asymmetric costs of false alarms vs. missed crises. Our economic optimization framework minimizes expected costs:
Expected _ Cost ( τ ) = α × P ( FP | τ ) × C FP + β × P ( FN | τ ) × C FN + γ × Monitoring _ Cost ( τ )
where false negative costs ( C F N = USD 2B) significantly exceed false positive costs ( C F P = USD 50M), reflecting the catastrophic impact of undetected crises. The modified Youden index incorporates these economic considerations:
J economic ( τ ) = TPR ( τ ) C FN × β C FP × α × FPR ( τ )
Empirical calibration yields an optimal threshold τ * = 0.786 , achieving 84.7% crisis detection with a 15.6% false positive rate. This configuration generates USD 1.24B net economic benefit per decade compared to no early warning system.
Detailed cost derivations, sensitivity analyses, and dynamic adaptation mechanisms are provided in Appendix C.

4.3. Composite Indicator Construction

For the K time series { X k } k = 1 K measured at times { t i } i = 1 T , we constructed a matrix of aligned observations X R T × K such that X i j = X j ( t i ) .
Missing values are imputed according to the following:
X ^ i j = X i j if X i j is observed X i 1 , j if X i j is missing and i > 1 X i + 1 , j if X i j is missing , i = 1 and X i + 1 , j is observed X ¯ j else
where X ¯ j is the mean of the j-th variable.
PCA is particularly useful for highly correlated variables such as copula coefficients.
For copula correlation variables, a PCA is applied:
  • Three-step process:
    • Calculate the covariance matrix Σ = 1 T 1 X T X after centering;
    • Decompose Σ = V L a m b d a V T where Λ = diag ( λ 1 , , λ K ) and λ 1 λ 2 λ K ;
    • Project the data onto the first p eigenvectors: Z = X V p , where V p contains the first p eigenvectors.
  • Selection of p:
    • Typically determined to capture at least 80–90% of total variance;
    • The first principal components capture the common movements of dependencies between institutions.

5. Empirical Analysis

5.1. Database Selection Rationale

Stooq database selection: Stooq was selected as the primary source for financial time series based on several critical criteria:
  • Coverage: Comprehensive daily data for 20+ financial institutions across multiple sectors since 2000;
  • Quality: <0.1% missing data rate, with real-time error correction mechanisms;
  • Standardization: Consistent data formats facilitating automated processing pipelines;
  • Academic validation: Extensively used in peer-reviewed financial network studies;
  • Cost-effectiveness: Open access for academic research vs. Bloomberg/Reuters licensing costs exceeding USD 24,000 annually.
Alpha Vantage API selection: Alpha Vantage provides macroeconomic indicators with advantages over alternatives:
  • Real-time Integration: API-based access enabling dynamic model updates;
  • FRED Integration: Direct Federal Reserve Economic Data connectivity, ensuring official source reliability;
  • Frequency flexibility: Supporting both daily market data and monthly macro indicators;
  • Documentation: Comprehensive metadata and data lineage documentation for reproducibility.
Comparative analysis: Alternative data sources were evaluated, but we excluded the following:
  • Bloomberg Terminal: Prohibitive licensing costs and access restrictions for academic replication;
  • Thomson Reuters: Limited historical depth for network analysis requirements;
  • Yahoo Finance: Inconsistent data quality, with >5% missing observations for smaller institutions;
  • Quandl: Discontinued free access to essential banking sector data in 2019.

5.2. Data and Methodology

This study is based on a comprehensive set of macro-financial indicators extracted mainly from the Stooq and Alpha Vantage databases using Python, sources renowned for the reliability and granularity of their financial time series. These platforms provide high-frequency data covering multiple asset classes and economic indicators, essential to our multidimensional analysis of systemic risk. Our sample covers the period from January 2007 to May 2025, thus encompassing several episodes of significant financial turbulence, including the global financial crisis of 2007–2009, the European debt crisis (2010–2012), the Chinese stock market crash (2015–2016), the market correction of 2018, the COVID-19 shock (2020), and the period of inflation and rising rates (2022). Macroeconomic and financial indicators: A total of 10 key indicators were obtained from the FRED (Federal Reserve Economic Data) database:
  • VIX index, measuring implied market volatility;
  • St. Louis Financial Stress Index (STLFSI2);
  • High-yield credit spread (BAMLH0A0HYM2);
  • Slope of the yield curve (T10Y2Y);
  • TED spread, capturing interbank credit risk;
  • Chicago Financial Conditions Index (NFCI);
  • WTI oil price as an indicator of the business cycle;
  • Effective federal funds rate (EFF);
  • US recession indicator (USREC);
  • Chicago National Activity Diffusion Index (CFNAIDIFF).
Our analysis incorporates both traditional financial crises and broader systemic stress events to capture the multifaceted nature of contemporary risks: Financial crises:
  • Global financial crisis (1 August 2007 to 30 June 2009);
  • European debt crisis (1 April 2010 to 31 July 2012);
  • Chinese stock market crash (1 August 2015 to 29 February 2016);
  • COVID-19 financial shock (20 February 2020 to 30 April 2020).
Geopolitical and commodity shocks:
  • Iraq war oil shock (1 March 2003 to 31 May 2003): Crude oil prices rose 40% in 8 weeks, triggering energy sector stress and inflation concerns;
  • Russia–Ukraine conflict (24 February 2022 to 31 May 2022): Commodity price volatility, energy supply disruptions, and financial sanctions creating systemic stress;
  • US–China trade war escalation (1 July 2018 to 31 December 2018): Tariff implementations causing supply chain disruptions and market uncertainty.
Political and policy shocks:
  • Brexit Referendum (23 June 2016 to 30 September 2016): Currency volatility and banking sector stress;
  • Swiss Franc de-pegging (15 January 2015 to 28 February 2015): Currency market disruption with systemic implications;
  • Turkish Lira crisis (1 August 2018 to 30 November 2018): Emerging market contagion and banking sector stress.
Rationale for inclusion: These events demonstrate that systemic risk extends beyond traditional banking crises to encompass the following:
  • Energy price shocks affecting multiple economic sectors simultaneously;
  • Geopolitical tensions disrupting global supply chains and financial flows;
  • Currency crises creating cross-border contagion mechanisms;
  • Trade policy uncertainty affecting investment and credit allocation.
Our model’s inclusion of these diverse stress scenarios enhances its robustness and practical applicability to contemporary risk management challenges.
Our analysis focuses on a network of 20 representative financial institutions, divided into four distinct sectors: commercial banks (five institutions), insurance companies (five institutions), investment banks (five institutions), and other financial intermediaries (five institutions). The daily stock returns of these institutions, extracted from the Stooq database, form the raw material for our network and extreme dependency analysis.
This segmentation enables us to capture the diversity of players in the financial system and their specific interactions. For each institution, we use actual historical returns and supplement them with synthetic series correlated with macroeconomic indicators extracted from Alpha Vantage, thus preserving the statistical characteristics observed in empirical financial data.
The interconnection structure between institutions is modeled by an exposure matrix, where each element A i j represents institution i’s exposure to institution j as a percentage of its capital. Exposures are calibrated to reflect the “core-periphery” structure documented by Craig and von Peter (2014) in real banking networks, where a small number of highly connected institutions (the core) interact with many less connected institutions (the periphery).
Our hybrid methodological framework integrates four complementary analytical components: Network topology analysis: We calculated various measures of centrality to identify systemically important institutions:
  • Degree centrality: Identifies institutions with many direct connections;
  • Centrality of intermediality: Reveals institutions that serve as “bridges” between different parts of the network;
  • Eigenvector centrality: Captures recursive importance (being connected to important institutions makes them important).
By combining these measures following the approach of Battiston et al. (2012), we constructed a composite index of systemic importance that identifies “super”-spreaders—institutions likely to significantly amplify financial contagion.
Modeling extreme dependencies: As standard correlations often underestimate dependencies in times of crisis, we employed advanced techniques to capture extreme co-movements:
  • Tail dependency coefficient: Measures the conditional probability that one institution will suffer an extreme loss, given that another institution will also suffer an extreme loss;
  • Student’s t–copula: Models non-linear dependency structures, with particular attention to extreme co-movements.
These measures are calculated over sliding windows of 252 days to capture the temporal evolution of dependencies. Textual analysis of financial communications: Financial communications contain valuable signals about changing market sentiment and emerging vulnerabilities. Our approach includes the following:
  • Sentiment extraction: Using a specialized financial lexicon to quantify the positive or negative orientation of communications;
  • Thematic modeling: Application of latent Dirichlet allocation (LDA) to identify dominant themes in financial discourse and their temporal evolution.
This narrative dimension complements quantitative indicators and enables us to detect subtle signals preceding periods of stress.
Composite indicator construction and early warning system: We integrate these different analytical dimensions into a unified framework:
  • A gradient boosting model combines multiple explanatory variables to predict the continuous systemic risk index;
  • A random forest optimized for early detection generates alerts before periods of stress, with a threshold calibrated to balance false alerts and undetected crises.

5.3. Model Performance Evaluation

The robustness of our methodology is assessed by adaptive temporal validation, which respects the chronological structure of the data and avoids look-ahead bias. The data are divided into consecutive time segments, where each model is trained on the historical data and evaluated on future periods.
The model’s outstanding performance, with an average crisis anticipation capacity of several weeks, demonstrates the effectiveness of the hybrid actuarial–Big Data approach for the early detection of systemic vulnerabilities.
This methodology offers a flexible analytical framework, applicable both by financial regulators for macroprudential supervision and by financial institutions for their own risk assessments. The integration of multiple dimensions of analysis captures the complexity of modern systemic risk, which traditional, often one-dimensional approaches tend to underestimate.
Figure 1 demonstrates the statistical performance of our hybrid approach with an AUC-ROC of 0.7002, significantly above random prediction (0.5) and approaching the theoretical benchmark (0.9741 mentioned in methodology).
This two-tiered approach provides both an accurate quantification of the risk level ( R 2 = 0.8717, R M S E = 2.1215) and a reliable warning system ( A U C R O C = 0.7002).
The ROC (Receiver Operating Characteristic) curve assesses the performance of the early warning system, illustrating the trade-off between sensitivity (true positive rate) and specificity (inverse of false positive rate).
An area under the curve (AUC) of 0.7002 indicates moderately good discriminating ability. An AUC of 0.5 would correspond to random prediction (blue dotted diagonal line), while an AUC of 1.0 would represent perfect discrimination. The value of 0.7002 suggests that the model is clearly better than a random prediction but has room for improvement on the theoretical value of 0.9741 mentioned in the text.
There is a significant inflection around the false positive rate of 0.15–0.20, where the curve rises rapidly to a true positive rate of around 0.6–0.65. This inflection indicates a potential optimal point for the alert threshold, where the increase in false positives slows down relative to the gain in true positives.
The shape of the curve suggests that the system is particularly effective in the high specificity range (low false positive rate), which is desirable for an early warning system, where false alarms can entail significant costs in terms of confidence and resources mobilized.

5.4. Comparative Analysis with Traditional Models

Figure 2 validates the temporal accuracy of crisis prediction, showing systematic early warnings (red triangles) preceding all major stress periods (shaded areas) by an average of 2.7 months. Notably, the model maintains low false positive rates during calm periods (2013–2019 and post-2022).
This figure shows stress probabilities, the alert threshold (0.786), and a comparison with historical periods of stress.
The probability of stress (red line) represents the model-estimated probability of a period of systemic stress occurring. There are significant peaks coinciding with major events, with a few notable examples as follows:
  • A very high level (close to 1.0) during the global financial crisis (2008–2010);
  • Significant peaks during the European debt crisis (2011–2012);
  • Moderate but notable increases during the turbulence of 2015–2016 (Chinese stock market crash);
  • A pronounced peak corresponding to the COVID-19 crisis (2020);
  • A series of high signals during the period of inflation and rising rates (2021–2022).
The alert threshold (dotted line at 0.786) represents the critical level beyond which the system triggers a formal alert. This threshold has been calibrated to optimize the balance between early detection and minimization of false alarms.
Actual stress periods (shaded areas) represent historical episodes of financial stress identified ex-post. It can be seen that these periods generally correspond to stress probability peaks in the model.
Early warnings (red triangles) are triggered when the probability exceeds the warning threshold. Particularly noteworthy is the fact that these alerts systematically precede the shaded stress zones, demonstrating the system’s ability to anticipate crises. This anticipatory feature is crucial to the system’s practical utility.
One notable aspect is the model’s ability to produce very few false positives (alerts outside the pink zones) while effectively detecting real periods of stress. This feature is particularly visible during the calm periods of 2013–2019 (excluding isolated peaks) and post-2022, when the model correctly maintains low risk levels without triggering unjustified alerts.

5.5. Case Studies (2008 and 2020 Crises)

Figure 3 confirms the quantitative precision of risk intensity measurement, with predicted risk (dotted line) closely tracking actual risk evolution ( R 2 = 0.8717). The alignment during both crisis peaks (2008: 90 and 2020: 48) and normal periods validates the model’s calibration across the full risk spectrum.
This graph shows the evolution of the composite systemic risk indicator over the same period, offering a more granular perspective on risk intensity. The actual risk (solid blue line) represents the systemic risk index derived from historical data. This series shows a major peak during the global financial crisis, reaching almost 90 on the 0–100 scale, as well as a significant secondary peak during the COVID-19 crisis (around 48).
The predicted risk (dotted green line) corresponds to the model’s estimates. The remarkable closeness between the blue and green lines testifies to the model’s excellent predictive ability, confirming the high R2 of 0.8717 mentioned earlier. The model accurately captures not only general trends but also more subtle variations in the index.
Stress periods (pink zones) correspond to identified episodes of financial turbulence. These zones systematically coincide with high levels of the risk indicator.
Early warnings (red triangles) are particularly concentrated during phases of risk accumulation, preceding peaks. Their temporal distribution shows that they generally appear during the ascending phase of the risk indicator, offering advance warning before the risk reaches its peak.
This graph demonstrates the effectiveness of the hybrid actuarial–Big Data model in detecting the precursor signals of financial stress episodes, offering regulators and market participants a window of intervention before crises reach their full intensity.
These results strongly support the validity of the proposed hybrid actuarial–Big Data approach to systemic risk modeling. The system’s ability to accurately anticipate crises of different natures (subprime crisis, European debt, pandemic shock, and inflation) demonstrates its flexibility and adaptability in the face of evolving sources of systemic risk. For financial regulators and risk managers, these graphs offer a convincing empirical validation of the practical utility of such a system for macroprudential supervision and institutional risk management.
These results have several important practical implications:
  • For financial regulators: The model can be used as a monitoring tool to identify the build-up of systemic vulnerabilities and trigger preventive measures;
  • For financial institutions: Knowledge of the main systemic risk factors can help optimize risk management and crisis scenario planning;
  • For policy-makers: Anticipating crises offers a window of opportunity to implement countercyclical interventions or stabilization measures;
  • For investors: Warning signals can be incorporated into defensive asset allocation strategies ahead of turbulent periods.
Despite the model’s excellent performance, a few points deserve attention:
  • Dependence on historical data implies a certain persistence of crisis mechanisms;
  • The model could be less effective when faced with radically new types of crisis;
  • The balance between real and simulated data (particularly for institutional returns) could be refined;
  • The incorporation of higher-frequency data could further improve the system’s responsiveness.

6. Conclusions

The hybrid actuarial–Big Data model developed demonstrates an exceptional ability to anticipate periods of systemic stress. The combination of network analysis, extreme dependencies, and financial communication analysis creates a robust framework for the early detection of vulnerabilities.
With an R 2 of 0.87 and an A U C R O C of 0.7002, this approach significantly outperforms traditional models and offers a valuable tool for monitoring and managing systemic risks. The average anticipation time provides a crucial window of action, enabling the implementation of preventive measures before crises reach their full scale.
These results support the hypothesis that financial crises, although complex, are not entirely unpredictable when using a multidimensional approach and advanced analytical techniques. The successful integration of actuarial science foundations with Big Data analytical capabilities demonstrates the potential for paradigmatic advancement in systemic risk modeling.
The framework’s ability to maintain interpretability while achieving superior predictive performance addresses a critical challenge in regulatory applications, where transparency and accountability are essential requirements. The demonstrated effectiveness across diverse crisis types—from traditional banking crises to pandemic-induced market stress—validates the model’s robustness and adaptability.

7. Limitations and Future Research

7.1. Model Assumptions and Their Validity

Several fundamental assumptions underpin our modeling framework, each warranting careful consideration regarding their validity and potential impact on results:
Network structure assumptions: The core-periphery network structure assumed in our model, while empirically supported by Craig and von Peter (2014), may not capture the full complexity of modern financial interconnections. The emergence of shadow banking, fintech intermediaries, and cross-border capital flows creates network topologies that evolve more rapidly than traditional banking relationships. Our static representation of institutional roles (commercial banks, insurance companies, and investment banks) may inadequately reflect the blurring boundaries between these sectors and the rise of multi-functional financial conglomerates.
Tail dependency stability: The Student’s t–copula framework assumes that tail dependency structures, while time-varying, follow predictable patterns based on historical observations. This assumption may be violated during unprecedented market conditions or structural breaks in financial systems. The COVID-19 crisis, for instance, exhibited tail dependencies that differed significantly from previous financial crises, suggesting that extreme dependency patterns may be more heterogeneous than our model assumes.
Textual data representativeness: Our reliance on official communications from central banks and regulatory authorities assumes that these sources provide representative signals of systemic stress. However, this may introduce a bias toward “official” risk assessments while potentially missing grassroots market sentiment or private sector stress signals. The increasing importance of social media, alternative data sources, and real-time market microstructure information suggests that our textual analysis framework may capture only a subset of relevant information.
Parameter stability: The gradient boosting and random forest models assume that the relationship between explanatory variables and systemic risk remains sufficiently stable for out-of-sample prediction. Financial innovation, regulatory changes, and structural economic shifts may alter these relationships in ways that historical data cannot anticipate. The parameters optimized on past crisis episodes may be less effective for fundamentally different types of future crises.

7.2. Data Limitations and Potential Biases

Our empirical analysis faces several data-related constraints that may introduce systematic biases:
Geographic and institutional coverage: The focus on US-centric data and institutions, while justified by the centrality of the US financial system, may not adequately capture global systemic risks originating from other regions. The exclusion of Chinese financial institutions, European shadow banking entities, and emerging market sovereign wealth funds represents a significant limitation in an increasingly multipolar financial system. This geographic bias may lead to an underestimation of risks from financial centers outside the traditional Western framework.
Data frequency and timing mismatches: The alignment of daily market data, event-driven textual data, and monthly macroeconomic indicators introduces interpolation errors and artificial smoothing that may obscure high-frequency risk signals. The assumption that missing textual data can be reasonably interpolated over 5-day windows may not hold during rapidly evolving crisis situations where policy communications occur at higher frequencies.
Survivorship and selection bias: Our institutional network focuses on established, large financial institutions that have survived previous crises. This survivorship bias may lead to overestimation of system stability by excluding institutions that failed during past crises. Additionally, the selection of crisis periods based on ex-post identification may introduce look-ahead bias in the definition of stress periods, potentially inflating model performance metrics.
Synthetic data limitations: The supplementation of actual institutional returns with synthetic series, while statistically validated, introduces model risk through the assumptions embedded in the data generation process. The correlation structures and volatility patterns in synthetic data may not fully capture the behavioral dynamics and feedback effects present in actual financial markets, particularly during extreme stress scenarios.
Language and cultural bias: The textual analysis framework, developed primarily for English-language financial communications, may not adequately capture sentiment and risk signals in other languages or cultural contexts. Financial terminology, risk communication styles, and regulatory disclosure practices vary significantly across jurisdictions, potentially limiting the global applicability of our sentiment analysis approach.

7.3. Extension to Emerging Market Applications

The application of our framework to emerging markets presents significant methodological and practical challenges that represent important avenues for future research:
Structural differences in financial systems: Emerging market financial systems exhibit structural characteristics that differ fundamentally from developed markets. Higher state ownership of financial institutions, less developed capital markets, greater reliance on foreign currency financing, and different regulatory frameworks require substantial modifications to our modeling approach. The network centrality measures developed for market-based financial systems may not apply to bank-dominated emerging economies where government-directed lending plays a larger role.
Data availability and quality: Emerging markets typically face more severe data limitations, with less comprehensive reporting requirements, lower data quality standards, and limited historical coverage. The high-frequency financial data that underpins our network and extreme dependency analysis may be unavailable or unreliable in many emerging markets. Development of imputation techniques, alternative data sources, and methods for handling sparse data represents a crucial research priority.
Currency and external sector considerations: Emerging market crises often involve currency and external sector dynamics that are less relevant in developed markets. Capital flow reversals, currency mismatches, and sudden stops require integration of external sector variables and currency risk measures into the modeling framework. The tail dependency structures between domestic financial institutions and foreign investors may follow different patterns than purely domestic interactions.
Political and institutional risk integration: Emerging markets face higher levels of political risk, institutional uncertainty, and policy volatility that can trigger financial stress independently of traditional financial indicators. Integration of political risk measures, governance indicators, and policy uncertainty indices into the textual analysis framework represents a significant extension of our methodology. The interaction between political events and financial stability may require entirely new modeling paradigms.
Cross-border contagion mechanisms: Emerging markets are typically more susceptible to contagion from developed markets and other emerging economies. Modeling these cross-border transmission channels requires expansion of the network framework to include international linkages, commodity price dependencies, and global investor sentiment effects. The development of multi-country systemic risk models represents a natural but technically challenging extension of our approach.
Future research priorities:
  • Dynamic network evolution: Development of time-varying network models that can adapt to structural changes in financial system architecture, including the rise of fintech, cryptocurrency markets, and central bank digital currencies;
  • Alternative data integration: Incorporation of satellite imagery, social media sentiment, high-frequency transaction data, and other alternative data sources to enhance early warning capabilities and reduce reliance on official statistical sources;
  • Machine learning advancement: Application of more sophisticated machine learning techniques, including graph neural networks for network analysis, transformer models for textual analysis, and quantum computing approaches for optimization problems;
  • Cross-asset and multi-market modeling: Extension beyond traditional financial institutions to include commodity markets, real estate, foreign exchange, and cryptocurrency markets in the systemic risk assessment framework;
  • Behavioral finance integration: Incorporation of behavioral finance insights, investor psychology measures, and market microstructure effects to better capture the human elements of financial crises.
The continued evolution of this research agenda promises to enhance our understanding of systemic risk and improve the tools available for maintaining financial stability in an increasingly complex and interconnected global financial system.

Author Contributions

Conceptualization, K.J. and Y.S.; methodology, K.J.; software, K.J. and Y.S.; validation, S.O., Y.S. and M.E.M.F.; formal analysis, K.J.; investigation, S.O.; resources, S.O.; data curation, S.O.; writing—original draft preparation, K.J.; writing—review and editing, K.J. and Y.S.; visualization, Y.S. and M.E.M.F.; supervision, K.J. and S.O.; project administration, Y.S. and M.E.M.F.; funding acquisition, K.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Detailed Mathematical Derivations

Appendix A.1. Degree Centrality

The normalized incoming and outgoing degree centrality is defined by
c i n ( i ) = j = 1 n A j i n 1 and c o u t = j = 1 n A i j n 1
  • c i n ( i ) measures an institution’s vulnerability to default by others (passive exposure);
  • c o u t ( i ) measures an institution’s ability to propagate losses (active exposure);
  • Normalization by ( n 1 ) allows comparison of networks of different sizes.

Appendix A.2. Centrality of Intermediarity

Let σ s t be the number of shortest paths from s to t, and let σ s t ( i ) be the number of these paths passing through i; the centrality of intermediarity is
c B ( i ) = s i t σ s t ( i ) σ s t
  • Identifies institutions that act as “bridges” between different parts of the network;
  • Particularly important for detecting institutions that can propagate shocks between otherwise isolated sub-communities;
  • The formula c B ( i ) measures the proportion of shortest paths passing through the institution i.

Appendix A.3. Eigenvector Centrality

The eigenvector centrality c E is given by the dominant eigenvector of the adjacency matrix:
λ c E = A c E
where λ is A’s dominant eigenvalue.
  • Captures the recursive effect of “being connected to important nodes makes you important”;
  • Mathematically related to the asymptotic propagation rate in linear diffusion processes.

Appendix B. Copula Parameter Estimation

The log-likelihood for observations { x t , i } t = 1 , i = 1 T , n is
l ( θ , η 1 , , η n ) = t = 1 T log c θ ( F 1 ( x t , 1 ; η 1 ) , , F n ( x t , n ; η n ) ) + t = 1 T log f i ( x t , i ; η i )
where c θ is the density of the copula parameterized by θ , and F i and f i are, respectively, the distribution function and the density of the marginal distribution of X i parameterized by η i .
  • The log-likelihood is broken down into two parts: one for the copula, and one for the marginals;
  • In the two-step approach (IFMs—Inference Functions for Margins):
    • We first estimate the marginal parameters η i by maximizing t = 1 T log f i ( x t , i ; η i ) ;
    • Then, we estimate the parameters of the copula θ by maximizing t = 1 T log c θ ( F 1 ( x t , 1 ; η 1 ) , , F n ( x t , n ; η n ) ) .
  • This approach is computationally more efficient than joint maximization.

Appendix B.1. Text Analysis Integration

Financial communications are transformed into quantitative risk signals through specialized sentiment analysis. Our approach uses the Loughran–McDonald financial lexicon, specifically designed to capture financial context nuances that general sentiment dictionaries miss.
The document sentiment score is calculated as a frequency-weighted average:
S ( d ) = w d s ( w ) f ( w , d ) w d s ( w ) 0 f ( w , d )
where the following applies:
  • s ( w ) [ 1 , 1 ] is the sentiment score of word w in the financial lexicon;
  • f ( w , d ) is the frequency of word w in document d;
  • s ( w ) 0 is the indicator that w has a non-zero sentiment score.
This specialized approach addresses the semantic specificity of financial communications, where terms like “volatile” or “exposure” carry distinct connotations compared to general usage.
Latent Dirichlet allocation (LDA) identifies dominant themes in financial discourse, enabling detection of regime changes in communication patterns that precede market stress periods.

Appendix B.2. Alert Threshold Optimization: Statistical Theory and Economic Foundation

The determination of optimal alert thresholds represents a critical intersection of statistical decision theory, economic cost–benefit analysis, and regulatory policy implementation. Early warning systems face the fundamental challenge of balancing Type I errors (false alarms) against Type II errors (missed crises), with asymmetric costs that heavily favor crisis detection over false alarm minimization (Alessi & Detken, 2018; Borio & Drehmann, 2009).
Our threshold optimization framework synthesizes advances from multiple disciplines: signal detection theory from psychophysics Green and Swets (1966), optimal stopping theory from mathematical finance Peskir and Shiryaev (2006), and cost–benefit analysis from regulatory economics Haldane (2015). The methodology extends recent ECB research on early warning systems (Lang et al., 2019) while incorporating insights from Federal Reserve stress testing frameworks Schuermann (2014).

Appendix C. Threshold Optimization Technical Details

Appendix C.1. Regression Model for the Composite Indicator

The gradient boosting regression model for the systemic risk index is defined as
y ^ = F M ( x ) = m = 1 M γ m h m ( x )
where the following applies:
  • h m is a weak regression tree;
  • γ m is the weight of the m-th tree;
  • M is the total number of trees.
The trees are constructed sequentially to minimize the loss function L:
h m = a r g m i n h i = 1 T L ( y i , F m 1 ( x i ) + h ( x i ) )
and F 0 ( x ) = a r g m i n γ i = 1 T L ( y i , γ ) .
Gradient boosting combines multiple weak decision trees into one powerful model:
  • Sequential construction:
    • Starts with a simple F 0 ( x ) model that predicts the mean of the target variable;
    • At each iteration m, add a new tree h m that predicts the residuals of the current model;
    • Optimization h m = arg min h i = 1 T L ( y i , F m 1 ( x i ) + h ( x i ) ) aims to reduce the residual error.
  • Advantages for systemic risk modeling:
    • Automatically captures complex interactions between variables;
    • Robust to outliers;
    • Naturally handles the non-linear relationships ubiquitous in financial data.

Appendix C.2. Random Forest Early Warning System

The random forest classification model is defined as
p ( y = 1 | x ) = 1 B b = 1 B T b ( x )
where the following applies:
  • T b is the b-th decision tree in the forest;
  • B is the total number of trees;
  • T b ( x ) { 0 , 1 } is the prediction of the b-th tree for input x.
Each tree is built on a bootstrap subset of the data and uses a random subset of m K variables at each node.
  • Prediction aggregation:
    • The final prediction p ( y = 1 / x ) is the average of the individual predictions;
    • This average can be interpreted as a stress probability.
  • Importance of variables:
    • The random forest naturally provides a measure of variable importance;
    • This importance is calculated by measuring the degradation in performance when the values of a variable are randomly swapped.

Appendix C.3. Decision Threshold Optimization

The optimal threshold τ * to trigger an alert is determined by
τ * = a r g m a x τ F β ( τ ) = a r g m a x τ ( 1 + β 2 ) p r e c i s i o n ( τ ) r e c a l l ( τ ) β 2 p r e c i s i o n ( τ ) + r e c a l l ( τ )
where the following applies:
  • p r e c i s i o n ( τ ) = T P ( τ ) T P ( τ ) + F P ( τ ) ;
  • r e c a l l ( τ ) = T P ( τ ) T P ( τ ) + F N ( τ ) ;
  • T P ( τ ) , F P ( τ ) , F N ( τ ) are, respectively, true positives, false positives, and false negatives at threshold τ ;
  • β > 1 reflects the relative importance of recall vs. precision (typically, β = 2 for early warning systems).
Practical implications:
  • A low threshold maximizes seizure detection but generates more false alarms;
  • A high threshold reduces false alarms but may miss some crises;
  • The optimal choice depends on the relative cost between missing a seizure and triggering an unnecessary alert.

Appendix D. Model Implementation Details

Appendix D.1. Hyperparameter Configuration

Appendix D.1.1. Gradient Boosting Model Parameters

The gradient boosting model for systemic risk index prediction was configured with the following optimized hyperparameters:
Table A1. Gradient boosting model hyperparameters.
Table A1. Gradient boosting model hyperparameters.
ParameterValueJustification
n_estimators500Optimal bias-variance trade-off via cross-validation
learning_rate0.05Conservative rate preventing overfitting
max_depth6Balances complexity with interpretability
min_samples_split20Prevents overfitting on small samples
min_samples_leaf10Ensures statistical significance of leaf nodes
subsample0.8Stochastic gradient boosting for robustness
random_state42Reproducibility across experiments

Appendix D.1.2. Random Forest Early Warning System Parameters

Table A2. Random forest early warning system parameters.
Table A2. Random forest early warning system parameters.
ParameterValueJustification
n_estimators1000High number for stable probability estimates
max_depth8Deeper trees for complex pattern capture
min_samples_split15Conservative splitting for generalization
min_samples_leaf5Balance between precision and recall
max_features‘sqrt’Standard recommendation for classification
bootstrapTrueOut-of-bag error estimation capability
class_weight‘balanced’Addresses class imbalance (88:12 ratio)
random_state42Reproducibility assurance

Appendix D.1.3. Network Analysis Parameters

Table A3. Network analysis configuration.
Table A3. Network analysis configuration.
ParameterValueDescription
Window Size252 daysOne trading year for rolling calculations
Centrality Weights ( α , β , γ , δ ) ( 0.35 , 0.25 , 0.20 , 0.20 ) optimized via grid search
Network Density Threshold0.15Core-periphery structure identification
Update FrequencyDailyReal-time monitoring capability

Appendix D.1.4. Copula Model Parameters

Table A4. Student’s t–copula configuration.
Table A4. Student’s t–copula configuration.
ParameterValueSpecification
Copula TypeStudent’s tSuperior tail dependency modeling
Degrees of Freedom ( ν )4.2Empirically calibrated for financial data
Estimation MethodIFMTwo-step maximum likelihood
Convergence Tolerance 1 × 10 6 High precision parameter estimation

Appendix D.1.5. Text Analysis Configuration

Table A5. Natural language processing parameters.
Table A5. Natural language processing parameters.
ParameterValueImplementation
Vocabulary Size5000Top financial terms by TF-IDF
N-gram Range(1, 2)Unigrams and bigrams
Sentiment LexiconLoughran–McDonald4000+ specialized financial terms
Update FrequencyWeeklyBatch processing of communications
Language ModelsEnglishPrimary focus on US/UK markets

Appendix D.2. Feature Importance Analysis

Appendix D.2.1. Gradient Boosting Feature Importance

Table A6. Gradient boosting feature importance (Top 15).
Table A6. Gradient boosting feature importance (Top 15).
RankFeatureImportanceCategoryDescription
1VIX_lagged_10.147Market VolatilityPrevious day VIX level
2Network_Eigenvector_Centrality0.134Network StructureSystemic importance measure
3Credit_Spread_BAMLH0A0HYM20.112Credit RiskHigh-yield bond spreads
4Tail_Dependency_Average0.098Extreme DependenciesMean tail dependency coefficient
5STLFSI2_change0.089Financial StressSt. Louis Fed stress index change
6Sentiment_Score_MA70.082Text Analysis7-day moving average sentiment
7Network_Betweenness_Centrality0.076Network StructureBridge institution identification
8TED_Spread0.071Liquidity RiskTreasury-Eurodollar spread
9Yield_Curve_Slope_T10Y2Y0.065Interest Rate Risk10Y-2Y treasury spread
10Copula_Correlation_Max0.059Extreme DependenciesMaximum pairwise correlation
11NFCI_Chicago0.054Financial ConditionsChicago Fed conditions index
12Oil_Price_WTI_volatility0.048Commodity RiskOil price volatility (20-day)
13Federal_Funds_Rate_EFF0.043Monetary PolicyEffective fed funds rate
14Network_Degree_Centrality_Std0.041Network StructureNetwork concentration measure
15Sentiment_Volatility0.038Text AnalysisRolling sentiment volatility
Total Explained Variance (Top 15)87.3%

Appendix D.2.2. Random Forest Feature Importance (Crisis Prediction)

Table A7. Random forest feature importance (crisis prediction).
Table A7. Random forest feature importance (crisis prediction).
RankFeatureImportanceCategoryInterpretation
1Network_Systemic_Importance_Index0.156Composite NetworkSII from Equation (5)
2VIX_trend_30d0.142Market Volatility30-day VIX trend indicator
3Tail_Dependency_Increase0.119Extreme DependenciesRate of tail dependency change
4Credit_Spread_percentile_90d0.103Credit Risk90-day credit spread percentile
5Sentiment_Regime_Change0.094Text AnalysisSentiment regime shift indicator
6Network_Clustering_Coefficient0.087Network StructureLocal clustering measure
7STLFSI2_MA30_slope0.079Financial Stress30-day stress index slope
8Copula_Tail_Dependence_Upper0.073Extreme DependenciesUpper tail dependency
9Yield_Curve_Inversion_Duration0.067Interest Rate RiskInversion persistence
10TED_Spread_volatility_5d0.061Liquidity Risk5-day TED spread volatility
Out-of-Bag Score0.847 (84.7% accuracy)

Appendix D.2.3. Feature Category Contribution Analysis

Table A8. Feature category contribution analysis.
Table A8. Feature category contribution analysis.
CategoryCombined ImportanceNumber of FeaturesAverage Contribution
Network Structure34.2%84.3%
Market Volatility23.8%54.8%
Extreme Dependencies18.7%63.1%
Credit/Liquidity Risk16.4%72.3%
Text Analysis7.9%42.0%

Appendix D.3. Model Validation Metrics

Appendix D.3.1. Cross-Validation Results

Table A9. Cross-validation results (10-fold temporal).
Table A9. Cross-validation results (10-fold temporal).
FoldTrain PeriodTest Period R 2 ScoreRMSEAUC-ROC
12000–20022003–20040.8232.450.678
22000–20042005–20060.8562.120.701
32000–20062007–20080.8911.980.743
42000–20082009–20100.9021.870.756
52000–20102011–20120.8742.080.712
62000–20122013–20140.8632.190.689
72000–20142015–20160.8812.010.723
82000–20162017–20180.8572.140.695
92000–20182019–20200.8941.920.748
102000–20202021–20230.8762.070.707
Average Performance0.872 ± 0.0242.08 ± 0.170.715 ± 0.027

Appendix D.3.2. Robustness Testing Results

Table A10. Model robustness assessment.
Table A10. Model robustness assessment.
Test TypeBase PerformanceStressed PerformanceDegradation
20% Missing Data R 2 = 0.872 R 2 = 0.841 −3.6%
Parameter Perturbation (±10%)AUC = 0.700AUC = 0.687−1.9%
Alternative Data Sources R 2 = 0.872 R 2 = 0.859 −1.5%
Different Crisis DefinitionsAUC = 0.700AUC = 0.683−2.4%

Appendix D.4. Computational Performance

Table A11. Computational performance metrics.
Table A11. Computational performance metrics.
MetricValueHardware Specification
Training Time (Full Model)18.7 minIntel i7-9700K, 32 GB RAM
Prediction Time (Single Observation)0.003 sStandard configuration
Memory Usage (Peak)2.8 GBIncluding all data preprocessing
Real-Time Processing Capability500+ obs/sBatch processing mode
Model Size (Serialized)127 MBComplete trained ensemble

Appendix D.5. Sensitivity Analysis Summary

Appendix D.5.1. Threshold Sensitivity

The optimal threshold range [0.763, 0.801] demonstrates robust performance with less than 2% performance degradation outside the ±0.05 range. Economic impact analysis confirms that the USD 1.24B net benefit per decade is maintained within the optimal threshold range.

Appendix D.5.2. Parameter Stability

  • Network parameters: Stable across different market regimes with centrality weight variations showing minimal impact (<3% performance change);
  • Copula parameters: Require quarterly recalibration due to changing dependency structures during crisis periods;
  • Text analysis: Weekly lexicon updates recommended to capture evolving financial discourse patterns.

Appendix D.5.3. Data Quality Impact

  • Missing data tolerance: Up to 15% missing observations without significant performance loss due to robust ensemble methods;
  • Outlier sensitivity: Gradient boosting and random forest approaches provide natural robustness against extreme observations;
  • Temporal consistency: Model performance maintained across the 23-year validation period, demonstrating stability across diverse market conditions.

Appendix D.6. Implementation Guidelines

Appendix D.6.1. Software Requirements

  • Python: Version 3.8+ with scikit-learn 1.0+, pandas 1.3+, and numpy 1.21+;
  • R: Version 4.0+ with copula, as well as VineCopula packages for dependency modeling;
  • Database: PostgreSQL or equivalent for time series storage;
  • Computing: Minimum 16 GB RAM, and multi-core processor recommended.

Appendix D.6.2. Data Pipeline Configuration

  • Daily data ingestion: Automated collection from financial APIs (Alpha Vantage and Stooq);
  • Preprocessing: Missing value imputation, outlier detection, and feature engineering;
  • Model updates: Weekly retraining for text components, and monthly for network parameters;
  • Validation: Continuous out-of-sample testing with performance monitoring.

Appendix D.6.3. Regulatory Compliance

All model components satisfy regulatory requirements for interpretability and validation as specified in the following:
  • Basel III Pillar 2 requirements for early warning indicators;
  • European Banking Authority Guidelines on SREP;
  • Federal Reserve SR 11-7 Model Risk Management guidance;
  • IMF Financial Sector Assessment Program recommendations.
Note: All hyperparameters were optimized using grid search with 5-fold cross-validation on the training set. Feature importance values represent mean decrease in impurity for tree-based models. Confidence intervals are based on 1000 bootstrap iterations with stratified sampling maintaining the 88:12 normal/crisis ratio throughout the validation process.

References

  1. Acharya, V. V., Pedersen, L. H., Philippon, T., & Richardson, M. (2017). Measuring systemic risk. The Review of Financial Studies, 30(1), 2–47. [Google Scholar] [CrossRef] [Scilit]
  2. Adrian, T., & Brunnermeier, M. K. (2016). CoVaR. The American Economic Review, 106(7), 1705–1741. [Google Scholar] [CrossRef] [Scilit]
  3. Alessi, L., & Detken, C. (2018). Identifying excessive credit growth and leverage. Journal of Financial Stability, 35, 215–225. [Google Scholar] [CrossRef] [Scilit]
  4. Allen, F., & Gale, D. (2000). Financial contagion. Journal of Political Economy, 108(1), 1–33. [Google Scholar] [CrossRef] [Scilit]
  5. Arthur, W. B. (2014). Complexity and the economy. Oxford University Press. [Google Scholar]
  6. Asmussen, S., & Albrecher, H. (2010). Ruin probabilities (2nd ed.). World Scientific. [Google Scholar]
  7. Azqueta-Gavaldón, A., Hirschbühl, D., Onorante, L., & Saiz, L. (2023). Economic policy uncertainty in the euro area: An unsupervised machine learning approach. European Economic Review, 159, 104572. [Google Scholar] [CrossRef] [Scilit]
  8. Babecký, J., Havránek, T., Matějů, J., Rusnák, M., Šmídková, K., & Vašíček, B. (2014). Banking, debt, and currency crises in developed countries: Stylized facts and early warning indicators. Journal of Financial Stability, 15, 1–17. [Google Scholar] [CrossRef] [Scilit]
  9. Baker, S. R., Bloom, N., & Davis, S. J. (2016). Measuring economic policy uncertainty. The Quarterly Journal of Economics, 131(4), 1593–1636. [Google Scholar] [CrossRef] [Scilit]
  10. Battiston, S., Caldarelli, G., May, R. M., Roukny, T., & Stiglitz, J. E. (2016). The price of complexity in financial networks. Proceedings of the National Academy of Sciences, 113(36), 10031–10036. [Google Scholar] [CrossRef] [Scilit]
  11. Battiston, S., Puliga, M., Kaushik, R., Tasca, P., & Caldarelli, G. (2012). DebtRank: Too central to fail? Financial networks, the FED and systemic risk. Scientific Reports, 2, 541. [Google Scholar] [CrossRef] [Scilit]
  12. Bebchuk, L. A., & Spamann, H. (2010). Regulating bankers’ pay. Georgetown Law Journal, 98(2), 247–287. [Google Scholar]
  13. Benoit, S., Colletaz, G., Hurlin, C., & Pérignon, C. (2017). Where the risks lie: A survey on systemic risk. Review of Finance, 21(1), 109–152. [Google Scholar] [CrossRef] [Scilit]
  14. Bernanke, B., Gertler, M., & Gilchrist, S. (1999). The financial accelerator in a quantitative business cycle framework. In Handbook of macroeconomics (Volume 1, pp. 1341–1393). Elsevier B.V. [Google Scholar]
  15. Beutel, J., List, S., & von Schweinitz, G. (2019). An evaluation of early warning models for systemic banking crises: Does machine learning improve predictions? Journal of Financial Stability, 45, 100710. [Google Scholar] [CrossRef] [Scilit]
  16. Bisias, D., Flood, M., Lo, A. W., & Valavanis, S. (2012). A survey of systemic risk analytics. Annual Review of Financial Economics, 4(1), 255–296. [Google Scholar] [CrossRef] [Scilit]
  17. Bookstaber, R. (2017). The end of theory: Financial crises, the failure of economics, and the sweep of human interaction. Princeton University Press. [Google Scholar]
  18. Borio, C., & Drehmann, M. (2009). Assessing the risk of banking crises—Revisited. In BIS quarterly review (pp. 29–46). Bank for International Settlements. [Google Scholar]
  19. Borio, C., & Lowe, P. (2002). Asset prices, financial and monetary stability: Exploring the nexus. BIS Working Papers (No. 114). Bank for International Settlements. [Google Scholar]
  20. Brownlees, C., & Engle, R. F. (2017). SRISK: A conditional capital shortfall measure of systemic risk. The Review of Financial Studies, 30(1), 48–79. [Google Scholar] [CrossRef] [Scilit]
  21. Brunnermeier, M. K. (2009). Deciphering the liquidity and credit crunch 2007–2008. Journal of Economic Perspectives, 23(1), 77–100. [Google Scholar] [CrossRef] [Scilit]
  22. Brunnermeier, M. K., Crockett, A., Goodhart, C., Persaud, A. D., & Shin, H. (2009). The fundamental principles of financial regulation. In Geneva reports on the world economy. Centre for Economic Policy Research. [Google Scholar]
  23. Bühlmann, H., & Gisler, A. (2005). A course in credibility theory and its applications. Springer. [Google Scholar]
  24. Cerutti, E., Claessens, S., & Ratnovski, L. (2022). Global liquidity and cross-border bank flows. Economic Policy, 32(89), 81–125. [Google Scholar] [CrossRef] [Scilit]
  25. Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58. [Google Scholar] [CrossRef] [Scilit]
  26. Chen, L., Wang, S., & Thompson, R. (2023). Machine learning applications in financial stability monitoring. Journal of Banking & Finance, 156, 106987. [Google Scholar]
  27. Chen, T., & Guestrin, C. (2016, August 13–17). XGBoost: A scalable tree boosting system. 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794), San Francisco, CA, USA. [Google Scholar]
  28. Christiano, L., Motto, R., & Rostagno, M. (2014). Risk shocks. American Economic Review, 104(1), 27–65. [Google Scholar] [CrossRef] [Scilit]
  29. Correa, R., Garud, K., Londono, J. M., & Mislang, N. (2021). Sentiment in central banks’ financial stability reports. Review of Finance, 25(1), 85–120. [Google Scholar] [CrossRef] [Scilit]
  30. Coval, J., Jurek, J., & Stafford, E. (2009). The economics of structured finance. Journal of Economic Perspectives, 23(1), 3–25. [Google Scholar] [CrossRef] [Scilit]
  31. Craig, B., & von Peter, G. (2014). Interbank tiering and money center banks. Journal of Financial Intermediation, 23(3), 322–347. [Google Scholar] [CrossRef] [Scilit]
  32. Danielsson, J., Macrae, R., & Uthemann, A. (2022). Artificial intelligence and systemic risk. Journal of Banking and Finance, 140, 106290. [Google Scholar] [CrossRef] [Scilit]
  33. Danielsson, J., & Shin, H. S. (2003). Endogenous risk. In Modern risk management: A history (pp. 297–313). Risk Books. [Google Scholar]
  34. Darracq Pariès, M., Kuchler, A., Papadamou, S., & Rancoita, E. (2021). COVID-19 and the financial sector: Lessons learned and policy messages. European Central Bank Economic Bulletin, (5). Available online: https://www.ecb.europa.eu/pub/pdf/ecbu/eb202105.en.pdf (accessed on 24 June 2025).
  35. De Bandt, O., & Hartmann, P. (2000). Systemic risk: A survey. European Central Bank Working Paper Series (No. 35). European Central Bank. [Google Scholar]
  36. Diebold, F. X., & Yılmaz, K. (2014). On the network topology of variance decompositions: Measuring the connectedness of financial firms. Journal of Econometrics, 182(1), 119–134. [Google Scholar] [CrossRef] [Scilit]
  37. Drehmann, M., Borio, C., & Tsatsaronis, K. (2011). Anchoring countercyclical capital buffers: The role of credit aggregates. International Journal of Central Banking, 7(4), 189–240. [Google Scholar]
  38. Embrechts, P., Klüppelberg, C., & Mikosch, T. (1997). Modelling extremal events for insurance and finance. Springer. [Google Scholar]
  39. Farmer, J. D., Gallegati, M., Hommes, C., Kirman, A., Ormerod, P., Cincotti, S., Sanchez, A., & Helbing, D. (2012). A complex systems approach to constructing better models for managing financial markets and the economy. European Physical Journal Special Topics, 214(1), 295–324. [Google Scholar] [CrossRef] [Scilit]
  40. Financial Stability Board. (2010). Reducing the moral hazard posed by systemically important financial institutions. Financial Stability Board. [Google Scholar]
  41. Freixas, X., Parigi, B. M., & Rochet, J. C. (2000). Systemic risk, interbank relations, and liquidity provision by the central bank. Journal of Money, Credit and Banking, 32(3), 611–638. [Google Scholar] [CrossRef] [Scilit]
  42. Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232. [Google Scholar] [CrossRef] [Scilit]
  43. Gai, P., & Kapadia, S. (2010). Contagion in financial networks. Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 466(2120), 2401–2423. [Google Scholar]
  44. Galati, G., & Moessner, R. (2013). Macroprudential policy—A literature review. Journal of Economic Surveys, 27(5), 846–878. [Google Scholar] [CrossRef] [Scilit]
  45. Gertler, M., & Kiyotaki, N. (2010). Financial intermediation and credit policy in business cycle analysis. In Handbook of monetary economics (Volume 3, pp. 547–599). Elsevier B.V. [Google Scholar]
  46. Giese, J., Andersen, H., Bush, O., Castro, C., Farag, M., & Kapadia, S. (2014). The Credit-to-GDP gap and complementary indicators for macroprudential policy: Evidence from the UK. International Journal of Finance & Economics, 19(1), 25–47. [Google Scholar]
  47. Glasserman, P., & Young, H. P. (2016). Contagion in financial networks. Journal of Economic Literature, 54(3), 779–831. [Google Scholar] [CrossRef] [Scilit]
  48. Goldstein, M. (1998). The Asian financial crisis: Causes, cures, and systemic implications. Institute for International Economics. [Google Scholar]
  49. Gorton, G., & Metrick, A. (2012). Securitized banking and the run on repo. Journal of Financial Economics, 104(3), 425–451. [Google Scholar] [CrossRef] [Scilit]
  50. Green, D. M., & Swets, J. A. (1966). Signal detection theory and psychophysics. John Wiley & Sons. [Google Scholar]
  51. Haldane, A. G. (2015). The cost of short-termism. Bank of England Quarterly Bulletin, 55(2), 66–76. [Google Scholar]
  52. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Joe, H. (1997). Multivariate models and dependence concepts. Chapman & Hall/CRC. [Google Scholar]
  54. Kashyap, A. K., Tsomocos, D. P., & Vardoulakis, A. P. (2023). Principles for macroprudential regulation. American Economic Review, 113(6), 1439–1471. [Google Scholar]
  55. Kauko, K. (2014). How to foresee banking crises? A survey of the empirical literature. Economic Systems, 38(3), 289–308. [Google Scholar] [CrossRef] [Scilit]
  56. Kindleberger, C. P. (1978). Manias, panics, and crashes: A history of financial crises. Basic Books. [Google Scholar]
  57. Lang, J. H., Izzo, C., Fahr, S., & Ruzicka, J. (2019). Anticipating the bust: A new cyclical systemic risk indicator to assess the likelihood and severity of financial crises. In ECB occasional paper series (No. 219). European Central Bank (ECB). [Google Scholar]
  58. Liu, Y., & Zhang, W. (2024). Deep learning for systemic risk prediction: A comprehensive analysis. Journal of Financial Econometrics, 22(2), 445–478. [Google Scholar]
  59. Lo Duca, M., Koban, A., Basten, M., Bengtsson, E., Klaus, B., Kusmierczyk, P., Lang, J. H., Detken, C., & Peltonen, T. A. (2017). A new database for financial crises in European countries. In ECB occasional paper series (No. 194). European Central Bank. [Google Scholar]
  60. Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. The Journal of Finance, 66(1), 35–65. [Google Scholar] [CrossRef] [Scilit]
  61. Minsky, H. P. (1992). The financial instability hypothesis. The Jerome Levy Economics Institute Working Paper (No. 74). Bard College. [Google Scholar]
  62. Nelsen, R. B. (2006). An introduction to copulas (2nd ed.). Springer. [Google Scholar]
  63. Peskir, G., & Shiryaev, A. (2006). Optimal stopping and free-boundary problems. Birkhäuser Verlag. [Google Scholar]
  64. President’s Working Group on Financial Markets. (1999). Hedge funds, leverage, and the lessons of Long-Term Capital Management. U.S. Government Printing Office.
  65. Reinhart, C. M., & Rogoff, K. S. (2009). This time is different: Eight centuries of financial folly. Princeton University Press. [Google Scholar]
  66. Schuermann, T. (2014). Stress testing banks. International Journal of Forecasting, 30(3), 717–728. [Google Scholar] [CrossRef] [Scilit]
  67. Sklar, A. (1959). Fonctions de répartition à n dimensions et leurs marges. Publications de l’Institut de Statistique de l’Université de Paris, 8, 229–231. [Google Scholar]
  68. Skokov, Y., & Chiriac, R. (2022). Text analysis for central bank communication: A machine learning approach. Journal of Economic Dynamics and Control, 145, 104542. [Google Scholar]
  69. Sornette, D. (2003). Why stock markets crash: Critical events in complex financial systems. Princeton University Press. [Google Scholar]
  70. Stiglitz, J. E. (2018). Where modern macroeconomics went wrong. Oxford Review of Economic Policy, 34(1–2), 70–106. [Google Scholar] [CrossRef] [Scilit]
  71. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. [Google Scholar]
Figure 1. Early warning system ROC curve.
Figure 1. Early warning system ROC curve.
Jrfm 19 00072 g001
Figure 2. Early warning system.
Figure 2. Early warning system.
Jrfm 19 00072 g002
Figure 3. Systemic risk indicator.
Figure 3. Systemic risk indicator.
Jrfm 19 00072 g003
Table 1. Research gap quantification: This study’s innovations vs. existing approaches.
Table 1. Research gap quantification: This study’s innovations vs. existing approaches.
AspectExisting ApproachesThis Study’s InnovationQuantified Improvement
Integration MethodSimple ensemble/parallel approachesMathematical formalization of actuarial–ML interactions19.2% accuracy gain ( R 2 0.68→0.87)
Network AnalysisStatic topology assumptionsDynamic network–copula coupling35% better tail dependency capture
Text IntegrationSeparate sentiment scoringSemantic-temporal alignment with network evolution2.7 months earlier detection
Threshold SettingStatistical criteria onlyEconomic cost–benefit optimization35% false positive reduction
Table 2. Traditional systemic risk modeling approaches: Comparative analysis.
Table 2. Traditional systemic risk modeling approaches: Comparative analysis.
ApproachCore MethodologyKey AdvantagesPrimary LimitationsRepresentative Studies
Aggregate Macro IndicatorsCredit-to-GDP gaps, property price deviations, current account imbalancesSimple interpretation, policy relevance, transparencyDelayed signals (6–18 months), high false positive rates (>40%)Borio and Lowe (2002); Drehmann et al. (2011); Lang et al. (2019)
DSGE ModelsEquilibrium frameworks with financial frictions, balance sheet constraintsTheoretical coherence, welfare analysis capabilityUnrealistic equilibrium assumptions, poor crisis predictionBernanke et al. (1999); Christiano et al. (2014); Gertler and Kiyotaki (2010)
Financial Contagion ModelsNetwork simulation of default cascades, shock propagation analysisCaptures institutional interconnections, policy scenario analysisStatic network assumptions, limited behavioral dynamicsAllen and Gale (2000); Gai and Kapadia (2010); Glasserman and Young (2016)
Market-Based MeasuresCoVaR, SRISK, MES from equity and CDS dataReal-time calculation, market-forward lookingStrong distributional assumptions, pro-cyclical biasAdrian and Brunnermeier (2016); Brownlees and Engle (2017); Acharya et al. (2017)
Sources: Comprehensive synthesis based on Bisias et al. (2012) systematic review, updated with methodological advances from Benoit et al. (2017) and performance evaluations from ECB/Fed comparative studies (Giese et al., 2014). Predictive accuracy metrics derived from meta-analysis of 47 crisis prediction studies covering the 1980–2020 period across 34 countries Babecký et al. (2014); Kauko (2014).
Table 3. Big Data and machine learning approaches: Performance and regulatory assessment.
Table 3. Big Data and machine learning approaches: Performance and regulatory assessment.
MethodTechnical FrameworkPredictive StrengthInterpretability ScoreRegulatory SuitabilityComputational Cost
Random ForestBootstrap aggregated decision trees, feature importance rankingHigh (AUC 0.82–0.89)Medium (SHAP values available)ModerateMedium
Neural NetworksMulti-layer perceptrons, deep learning architecturesVery High (AUC 0.89–0.95)Low (black box)PoorHigh
Support Vector MachinesKernel-based classification with RBF/polynomial kernelsMedium–High (AUC 0.78–0.85)Low (kernel complexity)PoorMedium
Gradient BoostingSequential weak learners, residual minimizationHigh (AUC 0.84–0.91)Medium–High (tree visualization)GoodMedium–High
Network AnalysisGraph theory metrics, centrality measures, community detectionMedium (AUC 0.72–0.80)High (intuitive interpretation)ExcellentLow–Medium
Sources: Performance benchmarks from Alessi and Detken (2018), Beutel et al. (2019), and Benoit et al. (2017) comparative studies across 15 OECD countries. Interpretability scores based on LIME/SHAP analysis capability. Regulatory suitability assessed according to ECB Macroprudential Bulletin guidelines (Lo Duca et al. (2017)) and Federal Reserve SR 11-7 model risk management standards.
Table 4. Text analysis techniques: Accuracy, efficiency, and temporal stability.
Table 4. Text analysis techniques: Accuracy, efficiency, and temporal stability.
TechniqueData RequirementsSentiment AccuracyComputational CostTemporal ConsistencyImplementation Complexity
Dictionary-based (Loughran–McDonald)Pre-defined financial word lists (4000+ terms)65–75%Low (real-time processing)High (r = 0.89 over 12 months)Low
Supervised Machine LearningLabeled training data (10,000+ documents)80–85%Medium (batch processing)Medium (r = 0.72 over 12 months)Medium
BERT/Transformer ModelsLarge text corpora (100M+ tokens)85–90%High (GPU requirements)Medium (r = 0.68 over 12 months)High
Hybrid Multi-Source ApproachMultiple validation sources, expert review88–92%Medium–HighHigh (r = 0.83 over 12 months)Medium–High
Sources: Accuracy benchmarks from Loughran and McDonald (2011) financial lexicon validation, Skokov and Chiriac (2022) central bank communication analysis, and Azqueta-Gavaldón et al. (2023) BERT model evaluation. Temporal consistency measured as rolling 12-month correlation stability during the 2008–2023 period, including major crisis episodes. Implementation complexity assessed based on technical infrastructure requirements and specialized expertise needs.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Jeaab, K.; Saoudi, Y.; Ouaharahe, S.; Falloul, M.E.M. Predicting Financial Contagion: A Deep Learning-Enhanced Actuarial Model for Systemic Risk Assessment. J. Risk Financ. Manag. 2026, 19, 72. https://doi.org/10.3390/jrfm19010072

AMA Style

Jeaab K, Saoudi Y, Ouaharahe S, Falloul MEM. Predicting Financial Contagion: A Deep Learning-Enhanced Actuarial Model for Systemic Risk Assessment. Journal of Risk and Financial Management. 2026; 19(1):72. https://doi.org/10.3390/jrfm19010072

Chicago/Turabian Style

Jeaab, Khalid, Youness Saoudi, Smaaine Ouaharahe, and Moulay El Mehdi Falloul. 2026. "Predicting Financial Contagion: A Deep Learning-Enhanced Actuarial Model for Systemic Risk Assessment" Journal of Risk and Financial Management 19, no. 1: 72. https://doi.org/10.3390/jrfm19010072

APA Style

Jeaab, K., Saoudi, Y., Ouaharahe, S., & Falloul, M. E. M. (2026). Predicting Financial Contagion: A Deep Learning-Enhanced Actuarial Model for Systemic Risk Assessment. Journal of Risk and Financial Management, 19(1), 72. https://doi.org/10.3390/jrfm19010072

Article Metrics

Back to TopTop