Next Article in Journal
Modeling Positive Seasonal Time Series with Dynamic Precision: The Generalized BPSARMA Model
Previous Article in Journal
Interactions Between Business Cycles, Financial Cycles and Monetary Policy in South Africa
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Learning Rare Events: Deep Learning Approaches to Extreme Price Prediction

School of Science and Technology, University of New England, Armidale, NSW 2350, Australia
*
Author to whom correspondence should be addressed.
Forecasting 2026, 8(3), 52; https://doi.org/10.3390/forecast8030052
Submission received: 17 May 2026 / Revised: 8 June 2026 / Accepted: 15 June 2026 / Published: 17 June 2026

Highlights

What are the main findings?
  • Deep learning approaches to price spike prediction are increasingly dominated by hybrid and transformer-based architectures, but forecasting performance is strongly influenced by spike definition, feature engineering, and class imbalance handling rather than architecture alone.
  • Only a small subset of reviewed studies explicitly formulate price spikes as a rare-event prediction problem, highlighting a major research gap in evaluation practices, spike-aware modelling, and robust real-world validation.
What are the implications of the main findings?
  • Future research should prioritise spike-aware problem formulation, imbalance mitigation, probabilistic forecasting, and interpretable evaluation metrics instead of focusing solely on increasingly complex neural architectures.
  • Accurate prediction of rare extreme price events has significant practical value for electricity markets, commodity trading, battery arbitrage, and financial risk management, particularly under increasing market volatility and renewable energy integration.

Abstract

Price spikes are rare but economically significant events observed across electricity, financial, commodity, and cryptocurrency markets. Their abrupt magnitude, heavy-tailed distributions, and severe class imbalance make them difficult to forecast using conventional time-series methods. This systematic literature review, conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) framework, synthesises recent deep learning approaches to forward-looking price-spike prediction and classification. Searches of Scopus, Web of Science, and IEEE Xplore identified studies published between 2020 and 2026. Following screening and full-text eligibility assessment of approximately 300 studies, only 20 met the inclusion criteria and were included in the final synthesis, comprising 19 peer-reviewed papers and one doctoral thesis. The review develops a structured taxonomy spanning spike definitions, task formulations, model architectures, input design, and evaluation practices. A central finding is that predictive performance is driven more by problem formulation, label construction, and evaluation design than by model architecture. While architectures have diversified to include recurrent networks, transformers, graph neural networks, and hybrid frameworks, improvements are often attributable to differences in how the prediction problem is defined rather than the models themselves. Key limitations stem from inconsistent spike definitions and insufficient treatment of class imbalance, leading to a misalignment between modelling objectives and evaluation practices, further exacerbated by the absence of standardised benchmarks. These issues hinder comparability and can lead to overstated model performance by masking poor detection of rare but economically critical spike events. The review therefore identifies clear directions for future research, including standardised spike labelling, adoption of rare-event-appropriate evaluation frameworks, and problem formulations that explicitly target extreme-event prediction.

1. Introduction

Extreme price movements, commonly referred to as price spikes, are rare but economically significant events, particularly in electricity markets where price dynamics are characterised by extreme volatility. Electricity has been shown to exhibit substantially higher volatility than other commodities, with annualised volatility approaching 300% in U.S. markets and exceeding 900% in the Australian National Electricity Market (NEM), compared to less than 100% for other energy commodities and below 20% for equity markets [1]. Moreover, a disproportionate share of average annual electricity prices can be attributed to spike events occurring in less than 1% of trading intervals, highlighting their economic significance despite their rarity.
These events are characterised by abrupt price changes, heavy-tailed distributions, structural breaks, and regime shifts [1,2,3,4,5]. Such properties violate assumptions of linearity and stationarity that underpin many traditional econometric and volatility modelling frameworks [6]. As a result, spike prediction represents a rare-event learning problem characterised by non-stationarity and severe class imbalance. In practice, such problems are typically addressed using approaches that either model the full price distribution or place greater emphasis on extreme observations, including volatility models, tail-risk frameworks, and data-driven methods with reweighting or resampling strategies to account for class imbalance.
Recent advances in deep learning have transformed time-series modelling. Architectures such as Long Short-Term Memory (LSTM) networks, Gated Recurrent Units (GRUs), Convolutional Neural Networks (CNNs), and attention-based models have demonstrated the ability to capture nonlinear temporal dependencies and high-dimensional multivariate relationships. Consequently, an expanding body of research applies deep learning techniques to spike prediction and classification tasks. However, substantial heterogeneity exists in spike definitions, label construction, input design, and evaluation metrics, which limits comparability across studies and complicates the interpretation of reported performance.
A systematic synthesis of this literature is therefore necessary to clarify methodological trends, identify areas of convergence and divergence in problem formulation, spike definition, modelling approaches, and evaluation practices, and highlight unresolved challenges in deep learning-based spike modelling. The analysis reveals that variation in predictive performance is largely explained by how the prediction problem is constructed and evaluated, rather than by the choice of neural architecture. This positioning also clarifies the scope of the present review as a focused subset of the broader forecasting literature.
This scope can be understood as a progressively narrowing subset of the broader forecasting problem, as illustrated in Figure 1, where rare event prediction forms a specialised domain within market forecasting, further constrained to time-series settings and, ultimately, to deep learning-based approaches.

1.1. Comparison with Existing Work

Existing review articles have examined electricity price forecasting, cryptocurrency forecasting, rare-event prediction, and transformer-based time-series forecasting [7,8,9,10,11,12,13,14]. However, reviews focused on forecasting applications primarily emphasise average price forecasting or general predictive accuracy rather than rare-event prediction, which has been extensively studied as a distinct machine learning problem in other domains [14]. When extreme prices are discussed, they are often treated as a secondary consideration within broader forecasting surveys [7,8,9,10,11,12]. This marginalisation leads to methodological choices that are poorly suited to rare events, including inappropriate loss functions and evaluation metrics, and ultimately limits progress in developing models that can reliably detect and anticipate economically critical price spikes. Table 1 summarises the scope and focus of these existing review studies and highlights the distinctions between prior work and the present review.
Reviews that address extreme prices, where considered, tend to emphasise traditional statistical and machine learning approaches such as regime-switching models, Support Vector Machines (SVMs), Auto-Regressive Integrated Moving Average (ARIMA), and Generalized Autoregressive Conditional Heteroskedasticity (GARCH) [9,10,11]. Deep learning is often presented as one technique among many within broader forecasting surveys [7,9,10,11]. Although dedicated reviews of deep learning forecasting methods have recently emerged [8], these reviews do not systematically categorise studies in terms of problem formulation, rare-event definition, architectural family, or evaluation design. In addition, most prior reviews focus on a single application domain, such as electricity markets [7,8,9,10,11] or cryptocurrency markets [12], limiting opportunities for cross-domain comparison and synthesis. This restricts the ability to assess the generalisability of forecasting methodologies and to isolate the impact of modelling choices on rare-event performance across market settings.
As a result, existing reviews provide limited clarity on whether reported improvements reflect genuine advances in spike prediction or are instead driven by differences in problem formulation, data preprocessing, or evaluation metrics. This lack of methodological alignment hinders cumulative progress and complicates comparison across studies.
Importantly, limited attention has been given to the implications of class imbalance, label construction, and metric selection for the interpretability and evaluation of rare-event prediction, despite broader surveys highlighting these challenges in general rare-event learning settings [14]. Furthermore, the rapid emergence of transformer-based architectures and hybrid deep learning frameworks is underrepresented in existing surveys, despite clear evidence of their rapid adoption across time-series forecasting tasks, including electricity price prediction [5,13,15].
While the limitations above relate to how prior reviews synthesise the field, similar issues are evident in the underlying empirical literature. A recurring pattern in the underlying literature is that studies claiming to address price spikes often do so only indirectly. Many papers retrieved through spike-related keywords ultimately focus on general price forecasting or related tasks, without explicitly formulating spike prediction as a dedicated rare-event learning problem. This creates a mismatch between how spike-related challenges are described and how they are operationalised and highlights the limited number of studies that treat spike prediction as a rare-event learning task in its own right.
In these studies, spike events are most commonly defined using threshold-based criteria, where prices exceeding a specified level are classified as extreme. However, standard modelling approaches optimise performance over the full price distribution. When such events are rare, optimisation is dominated by normal price behaviour, allowing models to achieve strong aggregate performance while failing to capture extreme events. This effect can be further amplified by preprocessing choices, such as normalisation or scaling, which compress extreme values and reduce their influence during training, particularly in models designed for general price prediction rather than rare-event detection. Accordingly, studies that do not explicitly model spike events as rare events are not considered to address rare-event prediction and are excluded from this review.
This review differs by explicitly treating spike prediction as a rare-event learning problem and by developing a structured taxonomy that spans spike definitions, model families, input configurations, and evaluation frameworks across market domains.

1.2. Objectives of This Review

The primary objective of this review is to systematically synthesise and categorise deep learning approaches for price spike prediction and classification across market domains. Specifically, the review aims to:
  • develop a structured taxonomy of spike definitions and modelling approaches
  • compare deep learning architecture families applied to spike tasks; examine input design choices, including univariate and multivariate configurations
  • analyse evaluation metrics used in rare-event contexts and assess their suitability and
  • identify methodological gaps and emerging research directions, particularly in probabilistic and spike-aware modelling frameworks.
This review goes beyond summarising existing work by identifying key gaps and outlining directions for the development of more rigorous, comparable, and theoretically grounded deep learning methods for extreme price forecasting.

2. Methods

This study follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) framework to ensure transparent and reproducible identification, screening, and synthesis of relevant literature [16]. The review process consisted of four stages: database search, screening of titles and abstracts, full-text eligibility assessment, and final inclusion. The study selection procedure is summarised using a PRISMA flow diagram presented in Figure 2.

2.1. Protocol and Registration

The review protocol was registered on the Open Science Framework (OSF) https://osf.io/mwfsx (accessed on 31 January 2026). The protocol specifies the search strategy, eligibility criteria, screening process, data extraction framework, and synthesis methodology. These components were defined a priori and are detailed in Section 2.2, Section 2.3, Section 2.4, Section 2.5, Section 2.6 and Section 2.7.

2.2. Search Strategy

2.2.1. Database Searches

A comprehensive search of Scopus, Web of Science, and IEEE Xplore was conducted, limited to publications from 2020 to 2026. This time window was selected to capture the period in which modern deep learning approaches, particularly attention-based architectures such as transformers, along with hybrid CNN–LSTM and sequence-to-sequence models, became widely adopted for time series forecasting [6,17,18,19,20,21,22,23]. Prior to 2020, most studies relied on traditional econometric methods or early machine learning techniques [24,25,26,27,28,29,30], which are not directly comparable to the current generation of models designed to capture non-linearity, long-range temporal dependencies, and complex cross-variable interactions [17,31]. Restricting the search to this period therefore ensures that the review reflects the contemporary methodological landscape and focuses on approaches relevant to current research and practical deployment in volatile market environments.
As shown in Table 2, equivalent search strategies were applied across databases using their respective broad search fields. In Scopus, all search terms were applied to the “Article title, Abstract, Keywords” field; in Web of Science, the “All Fields” option was used; and in IEEE Xplore, searches were conducted using “All metadata”. The lists of papers obtained from the Web of Science, Scopus and IEEE Xplore searches were compared to each other to filter duplicates. From the de-duplicated set, article titles were screened to assess relevance. Articles with ambiguous titles, as well as those clearly relevant, underwent full-text review.
Figure 2. PRISMA flow diagram of study selection. Records from Scopus, Web of Science, and IEEE Xplore were screened after filtering and deduplication, followed by full-text assessment and citation tracking. The final sample includes 20 studies (19 peer-reviewed, one doctoral thesis).
Figure 2. PRISMA flow diagram of study selection. Records from Scopus, Web of Science, and IEEE Xplore were screened after filtering and deduplication, followed by full-text assessment and citation tracking. The final sample includes 20 studies (19 peer-reviewed, one doctoral thesis).
Forecasting 08 00052 g002

2.2.2. Forward and Backward Citations

Backward snowballing was conducted by identifying all references cited within the current paper that were published from 2020 to Jan 2026 and that contained the terms “spike*” or “extreme price*” in the title. Forward snowballing was performed by using Google Scholar to compile the full list of publications citing the current paper. This citation list was then screened by searching title and abstract for the terms “price spike forecast*” and “price spike predict*” to identify relevant subsequent studies.

2.3. Eligibility Criteria and Study Selection

2.3.1. Inclusion Criteria

A study was included if it satisfied all the following:
  • Modelled forward-looking price spikes, extreme price events, flash crashes, or tail-risk exceedances
  • Uses time-series data with explicit temporal ordering
  • Applies deep learning models (e.g., LSTM, CNN, transformer)
  • Frames the task as forecasting, prediction, or prospective classification
  • Reports on quantitative evaluation metrics
  • Published in a peer-reviewed journal/conference or a doctoral thesis from a reputable institution
  • Written in English

2.3.2. Exclusion Criteria

Studies were excluded if they met any of the following:
  • Focus solely on post hoc anomaly detection without predictive intent
  • Use only statistical or shallow machine-learning models (e.g., ARIMA, GARCH, SVM) without deep learning
  • Do not involve price data (e.g., volatility indices only, order-book without prices)
  • Lack of sufficient methodological detail to reproduce the approach
  • Do not evaluate performance under class imbalance or extreme-event rarity
  • Are editorials, surveys or preprints without peer review
  • Duplicate studies reporting the same method and results

2.4. Study Screening Process

The screening process was conducted in two stages. First, titles and abstracts returned from the database searches were screened to identify potentially relevant studies. Articles that clearly did not address price spikes or extreme price events, did not apply deep learning models, or did not involve predictive modelling were excluded at this stage.
Studies that could not be confidently excluded based on title and abstract alone proceeded to full-text review. During full-text screening, articles were assessed against the predefined inclusion and exclusion criteria described in Section 2.3. Duplicate records across databases were removed prior to screening.
The screening process was conducted by a single reviewer. Eligibility criteria were defined a priori through protocol registration and applied consistently throughout the review process. Studies for which eligibility remained uncertain were discussed with supervisory team members, and final inclusion decisions were reached through consensus. Although this process was intended to minimise subjective bias and improve consistency, formal inter-rater reliability statistics were not calculated. This represents a limitation of the review and should be considered when interpreting inclusion decisions for borderline studies.

2.5. Data Extraction

For each included study, key methodological and contextual information was extracted using a structured data extraction template. The extracted fields included publication year, market domain, spike definition, task formulation (prediction, classification, or distributional modelling), deep learning architecture, input feature design, evaluation metrics, and reported outcomes.
Data extraction focused on characteristics necessary to construct the comparative taxonomy of spike modelling approaches described in Section 3. When studies reported multiple modelling configurations, the primary architecture and evaluation results discussed by the authors were recorded.

2.6. Study Quality Assessment

To assess methodological quality and potential sources of bias, each included study was evaluated against a set of criteria adapted for deep learning-based forecasting research. The assessment considered:
  • clarity of spike definition or extreme-event labelling procedure
  • appropriateness of evaluation metrics for rare-event prediction
  • treatment of class imbalance
  • transparency of model architecture and training procedure
  • presence of baseline comparisons or ablation analysis
  • use of temporally consistent train/test splits to avoid data leakage
The quality assessment was used to support interpretation of methodological trends and limitations across studies rather than to exclude papers from the review.

2.7. Synthesis Method

Due to substantial heterogeneity in spike definitions, modelling objectives, and evaluation metrics, a quantitative meta-analysis was not appropriate. Instead, a structured narrative synthesis was undertaken, designed as a technical comparison of modelling choices and predictive performance across studies.
Studies were systematically categorised by market domain, spike definition, task formulation, deep learning architecture, input design, and evaluation methodology. This taxonomy provided a consistent analytical framework to examine how design decisions influence spike prediction capability, enabling comparison of strengths and limitations across approaches and supporting broader conclusions that extend across application domains.

2.8. Common False Positives in the Screening Process

During the screening stage a substantial number of papers were initially retrieved because their titles or abstracts contained terms such as spike, extreme, or volatility. However, full-text review revealed that many of these studies did not address forward-looking spike prediction using deep learning and these excluded papers are listed in Table A1. Full-text analysis revealed five recurring categories of false positives, reflecting systematic mismatches between spike-related terminology and actual problem formulation.
First, many studies focus on general electricity price forecasting rather than spike prediction [32,33,34,35]. These papers typically predict day-ahead or real-time prices using deep learning models such as LSTM, CNN, or Transformer architectures but evaluate performance using regression metrics such as Root Mean Square Error (RMSE) or Mean Absolute Error (MAE). While such models may implicitly capture extreme price movements, they do not explicitly define or predict spike events.
Second, some studies present architectures that appear spike-aware but do not explicitly model spike occurrence [36,37,38]. These approaches incorporate mechanisms such as heteroscedastic loss functions, interval forecasting, or implicit long-tail modelling, which may capture extreme price movements within a continuous prediction framework. However, they lack a formal spike definition or classification objective and are not evaluated using metrics appropriate for rare-event detection. As a result, the ability of such models to identify spike events cannot be inferred from overall forecasting performance, and they cannot be meaningfully compared with approaches that directly optimise for spike detection.
Third, some papers frame the problem in terms of trading or optimisation rather than spike prediction [39,40]. In these studies, price forecasts are used as inputs to arbitrage or decision-making models, and performance is evaluated based on profitability rather than the accurate detection of rare extreme events.
Fourth, several studies focus on volatility modelling, anomaly detection, or retrospective analysis of extreme behaviour [41,42,43]. While these approaches are closely related to extreme price dynamics, they primarily aim to characterise, detect, or explain unusual behaviour in observed data rather than to forecast the occurrence of future spike events. In particular, many such methods operate retrospectively or contemporaneously, identifying anomalies or volatility regimes after they emerge, rather than learning to predict discrete spike events ahead of time. As a result, they do not address the forward-looking rare-event prediction task considered in this review and are not directly comparable with models explicitly designed to anticipate spike occurrence.
Finally, several studies employ probabilistic or distributional modelling techniques to capture tail behaviour without explicitly defining spike events [44,45]. These models estimate uncertainty or extreme quantiles but do not perform discrete spike prediction.
Because the objective of this review is to synthesise research on deep-learning approaches to forward-looking price-spike prediction, studies falling into the above categories were excluded during the eligibility assessment. It is important to emphasise that these exclusions do not reflect limitations in methodological quality or relevance of the underlying studies. Many of the excluded works represent state-of-the-art approaches to electricity price forecasting and make substantive contributions to the field. However, they are excluded on the basis of problem formulation rather than technical merit. Specifically, this review focuses on forward-looking predictions of rare, discrete spike events, whereas the majority of existing studies optimise expected error over the full price distribution. As these objectives are not equivalent, including such studies would confound the analysis by mixing fundamentally different learning problems under a single evaluation framework.

3. Results

This section classifies the included studies according to market domain, task formulation, spike definition, modelling architecture, input design, and evaluation strategy. Although the corpus spans multiple markets and modelling paradigms, several systematic patterns emerge across the literature.

3.1. Selection Results

A total of 300 records were identified through database searches across Scopus (102), Web of Science (111), and IEEE Xplore (87). After restricting results to publications from 2020 to 2026, 284 records remained. Duplicate records across databases were then removed, resulting in 191 unique studies for title and abstract screening.
Following initial screening, 124 articles were retained for full-text assessment. Studies were excluded at this stage if they did not address price spikes or extreme price events, did not employ deep learning methods, or lacked forward-looking predictive evaluation.
Twenty studies satisfied the inclusion criteria and were included in the final synthesis, comprising nineteen peer-reviewed papers and one doctoral thesis. The process is illustrated in the PRISMA flow diagram Figure 2. The final list of papers included in the review is shown in Table 3.

3.2. Market Domains

Electricity markets dominate the reviewed corpus. Most studies focus on wholesale electricity price forecasting in competitive or renewable-dominated environments [2,3,5,6,15,46,47,48,49,50,51,52,56,58]. These include day-ahead forecasting [46,48], multi-day forecasting [3], locational marginal price modelling [47,48,50], and extreme electricity price forecasting in real-time operational settings [5,6,52]. The geographical coverage spans several liberalised electricity markets, including the United States [5,47,48,49,50], Canada [3,15,58], China [6], and Australia’s National Electricity Market [2,51,52,54,56], reflecting the practical importance of spike prediction for system operators, market participants, and storage optimisation.
A smaller subset of studies extends beyond electricity markets as shown in Figure 3. Financial and cryptocurrency applications are represented in [4,55,57,59], while commodity price spike modelling appears in [53]. Despite differences in institutional structure and market drivers, the modelling approaches used in these domains closely resemble those used in electricity forecasting. Architectures, spike labelling methods, and evaluation metrics are largely transferable across domains, suggesting that spike prediction is best understood as a general rare-event forecasting problem rather than a phenomenon specific to electricity markets [14]. A notable distinction is that probabilistic and distributional approaches, including quantile-based forecasting and spike probability estimation, are increasingly explored across domains [5,50]. However, in electricity markets, spike prediction is more commonly operationalised through explicit thresholds or binary classification schemes [3,47,60], reflecting a stronger reliance on heuristic or market-defined spike definitions. These thresholds are typically heuristic or market-informed rather than statistically derived, introducing subjectivity into label construction and further limiting comparability across studies, as reflected in the absence of a consistent definition of spike magnitude or duration in the literature [58].

3.3. Problem Formulation and Spike Definition

Three primary task formulations can be identified across the literature. Although each seeks to characterise extreme price behaviour, they differ fundamentally in the learning target, optimisation objective, and the way spike events are defined. Throughout this section, Y t + h denotes the electricity price at forecast horizon h , X t denotes the set of predictor variables available at time t , Y ^ t + h denotes the corresponding model prediction, and τ denotes a spike threshold.
The first formulation frames spike prediction as a binary classification problem, where the objective is to predict whether a future price exceeds a predefined threshold. Spike events are typically defined using fixed price levels, statistical rules such as μ + k σ , or quantile-based criteria. Formally, the spike label is defined as
S t + h = 1 ( Y t + h > τ )
where 1 is the indicator function taking value 1 if the condition holds and 0 otherwise. The modelling task is to then estimate
P S t + h = 1 X t
This formulation appears in several electricity-focused studies [2,3,15,46,48,49,54], where class imbalance management and event detection reliability are central modelling concerns. Because P S t + h = 1 is typically small, optimisation and evaluation are oriented toward rare-event detection rather than overall accuracy, with metrics such as recall and F 1 -score used to assess performance on the positive class.
A second formulation adopts hybrid or regime-aware modelling frameworks that separate spike detection from price forecasting. In these architectures, a classification stage first estimates spike likelihood or latent market regime, followed by a regression model specialised for the predicted regime [6,46,47,48]. A simple representation is
R t + h = g X t , Y ^ t + h = f R t + h X t
where R t + h denotes the inferred market regime at horizon h   (e.g., normal or spike conditions), g is a mapping from inputs to regimes, and f R t + h denotes a regime-specific prediction function. Equivalently, the conditional expectation can be decomposed as
E   Y t + h X t = r P ( R t + h = r X t ) E   Y t + h X t , R t + h = r
Spike definitions in these models typically follow threshold-based criteria, as the classification stage requires explicit event labelling. These designs are motivated by the assumption that extreme price behaviour arises from mechanisms distinct from normal market dynamics.
The third formulation conceptualises spikes through distributional or probabilistic modelling rather than explicit classification. In these studies, models predict extreme price magnitudes, quantiles, volatility levels, or tail probabilities [4,5,50,51,52,57,59]. Instead of learning a binary label, the model estimates a conditional distribution
F Y X y X t = P Y t + h y X t
where F Y X is the conditional cumulative distribution function of future prices. Functionals of this distribution include the conditional quantile
Q α X t = F Y   |   X 1 α   |   X t
where Q α X t denotes the conditional α -quantile of the future price distribution and α 0 , 1 is typically chosen close to 1 to capture extreme prices. In this setting, spike occurrence is inferred indirectly through events such as P ( Y t + h > τ X t ) , rather than being explicitly labelled during training. This formulation is more common in financial and cryptocurrency applications [14] but increasingly appears in electricity forecasting as researchers adopt quantile-based and probabilistic prediction frameworks [3]. This formulation is motivated by the need to capture uncertainty and tail risk, particularly in financial and high-volatility markets where point forecasts are insufficient. By modelling the full conditional distribution, these approaches allow practitioners to estimate the likelihood and severity of extreme events rather than simply predicting their occurrence [14].
Taken together, these formulations correspond to different statistical objectives: event classification, regime-conditional prediction, and distributional inference. Their outputs are therefore not directly comparable, even when all are described as approaches to spike prediction. Across the literature, spike definitions vary substantially, with thresholds ranging from fixed price limits to statistical deviation rules, percentile-based criteria, and extreme value theory formulations, with no standardised approach emerging. These choices directly influence the prevalence of spike events, the degree of class imbalance, and the difficulty of the forecasting task. Consequently, differences in reported performance may reflect variations in problem formulation and label construction as much as genuine modelling improvements, limiting direct comparison across studies.
Importantly, task formulation and spike definition are not independent design choices. In classification-based approaches, the spike definition directly determines both the learning target and the prevalence of the positive class. More restrictive definitions generally result in rarer events and greater class imbalance, increasing the difficulty of model training and evaluation. In contrast, distributional models often avoid explicit labelling, instead inferring extreme events from predicted distributions. This interaction contributes to the observed heterogeneity across studies and complicates cross-study comparison. For practitioners, this implies that model selection should follow task formulation rather than preceding it, as different formulations require fundamentally different learning objectives and evaluation strategies.

3.4. Model Families

Recurrent neural networks, particularly LSTM-based architectures [61], remain the most widely used modelling backbone [3,6,46,47,48,49,50,52,56]. These models are designed for sequential data, updating a hidden state recursively so that predictions depend on both current inputs and past observations. LSTM variants introduce gating mechanisms that regulate information flow, enabling the model to capture longer-term temporal dependencies as shown in Figure 4. This makes them a natural starting point for spike prediction problems, where historical price dynamics, seasonality, and delayed system effects are important drivers. As a result, LSTM-based models are widely adopted as baseline architectures against which more specialised approaches are evaluated.
However, their dominance is not necessarily due to superior performance in rare-event settings. Rather, they provide a flexible and well-understood foundation that can be adapted to both regression and classification tasks. In practice, many studies report that vanilla LSTM models struggle to capture extreme events without additional modifications such as reweighting, feature engineering, or hybridization [46,48,50].
Hybrid architectures are therefore common. Several works employ two-stage spike-aware pipelines that combine classification and regression components [46,48,50], while others integrate deep learning with conventional machine learning algorithms such as support vector machines or gradient-boosted trees [48]. These designs are typically motivated by the assumption that extreme price events arise from market conditions distinct from those governing normal price variation, such as scarcity pricing, transmission congestion, or supply shocks [3,46,48]. By contrast, probabilistic forecasting approaches based on architectures such as Temporal Fusion Transformers model the full price distribution without explicitly separating spike events from normal price behaviour [17,45], and are therefore outside the scope of this review. These approaches may better capture gradual transitions between normal and extreme conditions and avoid error propagation between modelling stages. As a result, the relative effectiveness of hybrid versus unified approaches remains context-dependent and is not conclusively established in the current literature.
More recent studies explore alternative architectures to address specific limitations of RNN-based models. Graph neural networks (GNNs), as shown in Figure 5, are used to capture spatial dependencies between market regions or network nodes which is particularly relevant in interconnected systems where local events propagate across the network [2]. Transformer-based models have also gained attention due to their ability to model long-range temporal dependencies and complex feature interactions through attention mechanisms [15] as shown in Figure 6. These architectures are typically motivated by scalability and representation learning considerations. However, their advantages for spike prediction remain context-dependent, as larger datasets do not fundamentally address class imbalance, and rare events may still constitute only a small fraction of observations.
Although transformer architectures offer advantages in modelling long-range dependencies, they also introduce practical limitations. Self-attention mechanisms exhibit quadratic computational complexity with sequence length, increasing training cost and memory requirements. Furthermore, transformer performance is often highly sensitive to hyperparameter selection and may require substantially larger training datasets than recurrent architectures. In rare-event settings where spike observations are inherently limited, these characteristics may increase the risk of overfitting unless combined with appropriate regularisation, augmentation, or imbalance mitigation strategies.
In addition, a small number of recent studies extend beyond purely numerical inputs by leveraging architectures derived from large language models, particularly transformer-based frameworks originally developed for sequential text data and later adapted to time-series forecasting [51,52]. These approaches are often coupled with richer exogenous inputs or auxiliary forecasting modules, reflecting the observation that extreme price events cannot be reliably inferred from historical price series alone [5]. The underlying rationale is that price spikes are typically driven by external system conditions such as outages, congestion, or extreme weather, which are not fully observable within standard numerical time-series data [3]. This suggests that incorporating unstructured or contextual information sources has the potential to improve predictive performance in rare-event settings. Ensemble and neurosymbolic approaches further extend this idea by combining multiple models or integrating symbolic reasoning to stabilise predictions under high volatility [50,53], trading off simplicity for robustness and interpretability [53].
Despite this growing architectural diversity, empirical evidence suggests that improvements in spike prediction performance are often driven less by architectural novelty than by problem formulation, feature design, and imbalance mitigation strategies [3,5,15,46,58]. For practitioners, this implies that model selection alone is unlikely to be sufficient; careful consideration of how spikes are defined, represented, and weighted during training is often more critical than the choice of underlying architecture.

3.5. Inputs and Feature Engineering

Multivariate input structures dominate electricity forecasting models, incorporating historical prices, load forecasts, renewable generation, congestion indicators, and temporal features [3,5,6,46,47,48,50,52,56]. Many studies also introduce engineered variables designed to capture precursors of extreme price movements, including volatility measures, ramp indicators, and supply–demand imbalance metrics [46,48,56].
Graph-based approaches further incorporate structural information about network topology or regional interconnections [2,54], reflecting the spatial propagation of price spikes through transmission networks. Recent work also integrates contextual information sources, such as system operator notices or market announcements, using language models to extract signals associated with extreme market conditions [51,52]. These developments highlight the increasing role of external data sources in spike prediction research.

3.6. Evaluation Practices

Evaluation methodologies vary considerably across studies. Even among studies that explicitly target spike prediction, traditional regression metrics such as MAE and RMSE remain widely reported. These metrics evaluate average error across the full distribution of prices rather than focusing on extreme events. However, when spike events are rare ( P S = 1 1 ), expected loss is dominated by non-spike observations. Consequently, metrics defined over the full data distribution, such as
MAE = E [ Y Y ^ ]
RMSE = E [ ( Y Y ^ ) 2 ]
are primarily influenced by performance in the normal regime. This allows models to achieve low overall error while performing poorly on extreme events, thereby masking deficiencies in spike prediction.
To address this limitation, spike-focused studies more commonly employ precision, recall, and F1-score [2,13,17,22,24,25], which directly evaluate performance on the minority spike class. Letting TP, FP, TN, and FN denote true and false positives and negatives, these metrics are defined as
Precision = T P T P + F P ,   Recall = T P T P + F N ,   F 1 = 2 T P 2 T P + F P + F N
Unlike accuracy, which includes the true negative term and is therefore dominated by the majority non-spike class, these metrics exclude TN and depend only on TP, FP, and FN. Consequently, performance on rare spike events contributes directly to the evaluation, providing a more meaningful assessment under severe class imbalance.
Some studies also report ROC-AUC or balanced accuracy [2,54]; however, these metrics can be misleading in imbalanced settings. The ROC curve evaluates the trade-off between the true positive rate and false positive rate, where
TPR = T P T P + F N ,   FPR = F P F P + T N
When the number of non-spike observations is large, the denominator of FPR is dominated by TN, meaning that even a substantial number of false positives results in only a small increase in FPR. Consequently, models may achieve high ROC-AUC scores despite poor performance in identifying rare spike events, as strong classification of the dominant non-spike class masks deficiencies in rare-event detection.
Precision–recall analysis provides a more appropriate evaluation framework in this context, as it directly captures the trade-off between detecting rare events and avoiding false alarms. However, its use remains inconsistent across the literature. Decision threshold selection represents an additional but often overlooked source of performance variation. While metrics such as precision, recall, and F1-score evaluate classifier performance at a given operating point, the probability threshold used to classify spike events can substantially alter the balance between false positives and false negatives. In operational settings, these errors are rarely equally costly. For example, failing to anticipate a major electricity price spike may result in substantially greater economic loss than issuing a false alarm. Consequently, classification thresholds should ideally be selected according to application-specific cost functions rather than default probability cut-offs. However, relatively few studies explicitly justify threshold selection procedures or evaluate sensitivity to alternative operating points.
For models producing probabilistic spike forecasts, discrimination performance alone may be insufficient. A model that correctly ranks observations according to spike likelihood may still produce poorly calibrated probabilities. Calibration metrics such as the Brier Score and reliability analysis therefore provide complementary information by assessing whether predicted probabilities correspond to observed event frequencies. Despite their relevance for operational decision-making under uncertainty, calibration measures were rarely reported in the reviewed literature.
This misalignment between evaluation metrics and the rare-event nature of the task further complicates cross-study comparison and risks overstating model effectiveness. As a result, studies reporting strong performance under regression or ROC-based metrics may significantly overstate real-world spike detection capability. For practitioners, this implies that evaluation design is not a secondary consideration, but a primary determinant of whether model performance reflects true spike detection capability.

4. Discussion

The reviewed literature reveals that spike definition and task formulation exert a stronger influence on modelling outcomes than architectural choice. However, this finding should be interpreted in the context of the study’s inclusion criteria, which restrict the analysis to papers that explicitly formulate spike prediction as a rare-event learning task. Studies that treat extreme prices implicitly within general forecasting frameworks were excluded and may exhibit different relationships between model architecture and performance. Within the defined scope, differences in labelling, task formulation, and evaluation design consistently explain a substantial portion of the variation in reported results.
Across the corpus, electricity-focused studies predominantly adopt threshold-based definitions of extreme prices [2,3,15,46,48,49,56], transforming spike prediction into an imbalanced binary classification problem. In contrast, several studies conceptualise extreme prices through probabilistic or distributional frameworks [4,5,50,51,52,57,59], where spikes emerge as tail events within a continuous predictive distribution. These conceptual differences produce fundamentally different modelling pipelines, feature representations, and evaluation strategies. Consequently, comparisons across architectures are often confounded by differences in problem formulation rather than reflecting genuine improvements in predictive capability.
This observation can be formalised through the decomposition of expected risk. When spike events are defined as S = 1 ( Y > τ ) , the expected loss can be expressed as:
R f = P S = 0 E l S = 0 + P S = 1 E l S = 1
Given that P S = 1 1 , optimisation is dominated by non-spike observations. As a result, models trained under standard empirical risk minimisation are incentivised to minimise error in the normal regime, even when the objective of interest is accurate spike detection. This provides a theoretical explanation for the widespread empirical finding that models can achieve strong aggregate performance while systematically underperforming on extreme events.
A recurrent structural pattern in electricity forecasting studies is the use of hybrid or regime-aware modelling frameworks [46,48,50]. These approaches explicitly recognise that extreme price behaviour arises from market mechanisms distinct from those governing normal price fluctuations. Electricity markets exhibit nonlinear dynamics driven by scarcity pricing, transmission congestion, generator outages, and renewable intermittency. As a result, extreme price regimes are often characterised by abrupt structural changes rather than incremental deviations from normal conditions. Two-stage architectures attempt to capture this phenomenon by separating spike detection from conditional price forecasting. Formally, this corresponds to introducing a latent regime variable R , such that:
E Y X = r P R = r X E Y X , R = r
While such decompositions often improve spike recall and F1-score, they introduce a key limitation: errors in regime classification propagate downstream, constraining the performance of the regression stage and limiting recovery of misclassified spike events.
Class imbalance emerges as a central and unifying challenge across all formulations. Extreme prices typically represent a very small fraction of total observations, particularly in high-frequency electricity markets [58]. As a result, standard loss functions and aggregate regression metrics are often poorly aligned with the rare-event nature of the task. Many studies therefore reformulate the problem through explicit spike labelling, regime separation, hierarchical forecasting pipelines, or separate spike-classification stages to increase sensitivity to extreme events [6,15,46,47]. These approaches modify the effective contribution of rare events to the optimisation objective and frequently yield larger performance gains than architectural changes. This suggests that spike prediction is fundamentally a statistical rarity problem rather than a purely representational learning problem.
Architectural innovation nonetheless plays a role in capturing structural properties of electricity markets. Graph-based models [2,54] incorporate spatial relationships between network nodes, reflecting the propagation of price signals through transmission systems. Transformer-based models [15,55] enable the modelling of long-range temporal dependencies and complex feature interactions. However, the reviewed literature has yet to demonstrate consistent improvements in rare-event prediction relative to well-tuned recurrent baselines. In many cases, performance gains instead arise from feature engineering and data preprocessing [15,50,54,62]. Common feature engineering techniques include the construction of lagged variables, rolling statistics (e.g., moving averages, volatility measures, and quantiles), and domain-specific indicators such as demand–supply imbalances, reserve margins, and interconnector flows [15,50,54,62]. These features provide the model with explicit signals related to system stress conditions that are often associated with price spikes.
Data preprocessing methods further improve performance by addressing distributional and imbalance challenges inherent in spike prediction. These include scaling and normalisation to stabilise training, outlier handling or clipping to reduce the influence of extreme noise, and resampling or reweighting strategies to mitigate class imbalance [37,49,51]. In addition, some studies apply clustering or segmentation to isolate distinct operating regimes [15] or incorporate external data sources such as weather forecasts and market notices to capture exogenous drivers of extreme events [52,54]. Collectively, these techniques improve the signal-to-noise ratio and ensure that rare but informative spike events contribute more effectively to model training.
Recent work incorporating large language model augmentation [51,52] highlights a promising direction for improving spike prediction. These approaches typically integrate unstructured textual data, such as news reports, outage announcements, and weather forecasts, with structured numerical inputs using embedding techniques or multimodal architectures. For example, textual data can be encoded using transformer-based language models and fused with time series features within hybrid deep learning pipelines, enabling models to capture contextual signals that are not present in historical price data alone. This is particularly relevant in electricity markets, where extreme price events are often triggered by exogenous factors such as generation outages, transmission constraints, and extreme weather conditions [3].
While these methods are conceptually well motivated and represent an important step toward incorporating real world context into forecasting models, they remain at an early stage of development. Current studies are limited in scale, often rely on simplified data integration strategies, and lack systematic evaluation under realistic market conditions. In particular, the interaction between textual signals and rare event prediction remains underexplored, and it is not yet clear how much incremental predictive value such data provides relative to well-engineered numerical features.
Despite these advances in model inputs and architectures, evaluation practices represent a significant source of distortion in reported results. Many studies continue to rely on regression metrics such as MAE and RMSE [2,3,5,6,15,46,47,48,49,50,51,52,54,56,58], which primarily reflect performance on normal price observations and provide limited insight into spike detection capability. As a result, models can achieve low aggregate errors while systematically failing to capture extreme events. Similarly, ROC-AUC can remain high under severe class imbalance, as performance is dominated by the majority class, masking poor detection of rare spikes. This creates a disconnect between reported performance and operational usefulness, particularly in applications where the primary objective is to identify extreme price events. Precision–recall metrics offer a more appropriate evaluation framework by focusing on performance within the minority class, yet their use remains inconsistent across the literature.
The inconsistency of spike definitions further complicates cross-study comparison. Fixed thresholds, statistical deviation rules, percentile-based definitions, and tail-risk formulations coexist within the literature [53]. Because the frequency and severity of labelled spike events depend directly on the chosen definition, performance metrics cannot be meaningfully compared across studies without accounting for differences in labelling methodology. This lack of standardisation represents a fundamental limitation in current research.
Another structural limitation is the limited evaluation of model robustness under regime shifts. For example, electricity markets evolve over time due to renewable integration, regulatory changes, and shifting demand patterns [5,47,50]. Models trained on historical data may therefore fail to generalise under new conditions. Few studies explicitly evaluate performance across different market regimes or extended out-of-sample periods, despite several authors noting that electricity markets evolve over time due to structural and operational changes [15,63], limiting confidence in real-world applicability.
Taken together, the literature indicates that advances in spike prediction are driven primarily by methodological alignment between task formulation, spike definition, training objective, and evaluation strategy, rather than by neural architecture alone. Misalignment between these components leads to systematically biased performance estimates and obscures the true capability of proposed models.

4.1. Relevance of Methods to Other Domains

Spike prediction shares structural similarities with anomaly detection and extreme-event forecasting problems across multiple domains, including financial markets and multivariate spatiotemporal systems, where rare, high-impact events arise within nonlinear and non-stationary data-generating processes [50,64]. In each case, the central challenge is to identify these events reliably despite noisy observations, complex temporal dependencies, and evolving system dynamics that often violate the assumptions of traditional forecasting approaches [3].
The methodological components synthesised in this review, including threshold-based and anomaly-oriented labelling strategies, probabilistic and distributional modelling, multivariate feature integration, and hybrid or ensemble learning architectures, have been widely applied across domains characterised by rare-event dynamics [53,60]. In particular, formulating spike prediction as a classification or anomaly-detection problem reflects a broader paradigm in which extreme events are modelled as deviations from learned normal behaviour rather than as part of a continuous regression task [53].
Recent advances in deep learning further reinforce this cross-domain applicability. Attention-based architectures, including transformer models, have demonstrated strong capability in capturing long-range temporal dependencies and complex feature interactions in both electricity markets and high-frequency financial systems [15,55]. Empirical studies in real-world electricity markets have shown that such architectures can outperform traditional forecasting approaches, particularly under volatile conditions, highlighting their suitability for modelling rare and extreme price movements [15,63].
Across these domains, a consistent challenge is the presence of severe class imbalance and heavy-tailed distributions, which necessitates the use of specialised evaluation frameworks such as precision–recall metrics and probabilistic forecasting techniques. Approaches designed to address these challenges in electricity markets are therefore directly transferable to other rare-event prediction settings, including financial volatility modelling and risk-sensitive decision systems [3,57]. However, the literature lacks standardised benchmarks and datasets specifically designed for spike prediction, with most studies relying on proprietary or market-specific data and inconsistent spike definitions. This limits reproducibility and makes it difficult to compare model performance across studies.
By examining spike modelling through this focused methodological lens, this review contributes not only to the electricity price forecasting literature but also to the wider field of deep learning for extreme-event prediction. The synthesis highlights that several core challenges, including non-stationarity, data sparsity, and regime-dependent behaviour, also arise in other rare-event prediction settings, suggesting potential for methodological transfer. While the scope is deliberately narrow, focusing on studies that explicitly formulate spike prediction as a rare-event learning task, this enables a more rigorous comparison of modelling approaches and isolates the impact of problem formulation, labelling, and evaluation choices that are often obscured in broader surveys.

4.2. Implications for Practice and Future Research

The synthesis of findings suggests that progress in price spike prediction will depend on improved methodological alignment rather than continued architectural innovation. Based on the observed limitations, several recommendations emerge.
First, greater consistency in spike definition is required. Future studies should adopt either statistically grounded approaches, such as percentile-based or extreme value theory formulations, or clearly justify threshold selection in relation to market structure. Reporting multiple spike definitions where feasible would further improve robustness and comparability.
Second, evaluation practices must be aligned with the rare-event nature of the task. Precision–recall-based metrics, including F1-score and PR-AUC, should be prioritised over metrics dominated by majority-class performance. Studies should also report calibration measures or probabilistic scoring rules where applicable, to assess the reliability of predicted spike probabilities.
Third, baseline selection requires standardisation. Comparisons against simple but relevant baselines, such as persistence models, threshold heuristics, or classical statistical approaches, should be consistently included to contextualise performance gains. Without such baselines, it remains difficult to assess whether complex models provide meaningful improvements.
Fourth, class imbalance should be treated as a core modelling consideration rather than a secondary adjustment. Techniques such as cost-sensitive learning, resampling strategies, and threshold optimisation should be explicitly reported and evaluated, given their demonstrated impact on performance.
Fifth, robustness and generalisability require greater attention. Future studies should evaluate models across different market regimes, extended out-of-sample periods, and varying levels of volatility to ensure that performance is not confined to specific historical conditions.
Finally, the convergence of modelling approaches across electricity, financial, and commodity markets suggests an opportunity for cross-domain benchmarking. Establishing shared datasets or evaluation protocols would enable more systematic comparison and accelerate the development of transferable rare-event forecasting methods.
Table 4 presents a practitioner-oriented framework derived from the findings of this review. Rather than beginning with model architecture selection, the framework emphasises the importance of first defining the forecasting objective, data characteristics, uncertainty requirements, and operational constraints. Across the reviewed literature, forecasting performance was frequently influenced more by problem formulation, spike definition, feature design, and evaluation strategy than by architectural novelty alone. Consequently, selecting an appropriate forecasting formulation requires consideration of factors such as spike label availability, event rarity, uncertainty requirements, and the intended operational application. This perspective reinforces the central finding of the review that methodological alignment between forecasting objectives, data representation, and evaluation design is often more important than the choice of underlying deep learning architecture.

5. Conclusions

This review synthesised recent deep learning approaches to price spike prediction, revealing substantial heterogeneity in how the problem is formulated, labelled, and evaluated. While a wide range of architectures have been proposed, including recurrent, transformer-based, graph-based, and hybrid models, a central finding is that predictive performance is driven less by architectural choice than by upstream decisions regarding task formulation, spike definition, and evaluation methodology.
Across the literature, spike prediction is variously framed as a classification, hybrid, or distributional task, with no consistent alignment between formulation and application objectives. Similarly, spike definitions range from fixed thresholds to quantile-based and tail-risk approaches, directly influencing class imbalance, learning difficulty, and reported outcomes. This lack of standardisation complicates cross-study comparison and obscures the true sources of performance gains. In many cases, improvements attributed to model design instead reflect differences in how the prediction problem itself is constructed.
Evaluation practices further amplify this issue. The continued use of regression metrics and ROC-based measures in highly imbalanced settings can mask poor performance on rare events, while precision–recall metrics, though more appropriate, are applied inconsistently. As a result, reported model performance may not accurately reflect real-world spike detection capability, limiting the practical utility of many approaches.
Despite these contributions, a limitation of this review is the concentration of included studies in electricity markets, which likely reflects both data availability and terminology used in database searches. As a result, relevant work in other domains, such as cloud systems, telecommunications, or broader demand forecasting, may be underrepresented. This highlights the need for future reviews to adopt broader search strategies and cross-domain terminology when examining rare-event prediction problems.
For practitioners, these findings suggest that effective spike prediction requires careful alignment between problem formulation, feature design, and evaluation strategy, rather than reliance on increasingly complex model architectures. Techniques such as regime separation, targeted feature engineering, and imbalance-aware training and evaluation are often more impactful than architectural innovation alone. For researchers, the results highlight the need for greater methodological clarity and standardisation, particularly in the definition and evaluation of extreme events.
Future work should prioritise the development of consistent benchmarking frameworks, explicit treatment of rare-event learning challenges, and the integration of exogenous information sources that may drive extreme behaviour. A major limitation of the current literature is the absence of standardised benchmark datasets and evaluation protocols for price spike forecasting. Studies employ different markets, forecasting horizons, spike definitions, train-test splits, and evaluation metrics, making direct comparison of reported results difficult. Consequently, apparent performance improvements may reflect differences in dataset construction or problem formulation rather than genuine methodological advances.
Progress in the field would benefit from publicly available benchmark datasets, transparent spike labelling procedures, standardised train-test partitions, and consistent reporting of evaluation metrics such as precision, recall, F1-score, PR-AUC, and calibration measures. Reproducibility would also be strengthened through open-source baseline implementations and evaluation protocols incorporating rolling-origin validation and regime-shift testing, facilitating more rigorous comparison of forecasting approaches across electricity, commodity, cryptocurrency, and financial markets.

Author Contributions

Conceptualization, M.S., A.J.S. and F.H.; methodology, M.S., A.J.S. and F.H.; validation, A.J.S. and F.H.; formal analysis, M.S.; investigation, M.S.; resources, M.S., A.J.S. and F.H.; data curation, M.S.; writing—original draft preparation, M.S.; writing—review and editing, M.S., A.J.S. and F.H.; visualization, M.S.; supervision, A.J.S. and F.H.; project administration, M.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No additional data was created in this study.

Acknowledgments

This research was supported by the Commonwealth through an Australian Government Research Training Program Scholarship DOI: https://doi.org/10.82133/C42F-K220 (accessed on 20 January 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ANNArtificial Neural Network
ARIMAAutoregressive Integrated Moving Average
AUCArea Under Curve
CNNConvolutional Neural Network
EPFElectricity Price Forecasting
EVTExtreme Value Theory
GARCHGeneralised Autoregressive Conditional Heteroskedasticity
GNNGraph Neural Network
GRUGated Recurrent Unit
LSTMLong Short-Term Memory
LLMLarge Language Model
MAEMean Absolute Error
MAPEMean Absolute Percentage Error
MLMachine Learning
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
RMSERoot Mean Square Error
ROCReceiver Operating Characteristic
TFTTemporal Fusion Transformer

Appendix A

Table A1. List of papers considered but excluded in this study.
Table A1. List of papers considered but excluded in this study.
AuthorsTitleYearSource
X. Yang; P. Sha; J. Shen; J. Zhang; X. Wu; Y. LiA Novel Approach to Electricity Price Prediction: LDS-GRU Algorithm for Long-Tail Effects2024IEEE
Sawangtong, P.; Taleghani, R.; Taghipour, M.A novel framework for pricing electricity derivatives: Integrating stochastic models and machine learning2026Scopus
Karasu, S.A Novel Hybrid Model Using Demand Concentration Curves, Chaotic AFDB-SFS Algorithm and Bi-LSTM Networks for Heating Oil Price Prediction2025Scopus
V. Ramana K; M. KathiravanA Novel XG-Boost Approach for Electricity Load Balancing and Demand Prediction2023IEEE
Shao, Z.; Yang, Y.; Zheng, Q.; Zhou, K.; Liu, C.; Yang, S.A pattern classification methodology for interval forecasts of short-term electricity prices based on hybrid deep neural networks: A comparative analysis2022Scopus
Shi, W.; Feng Wang, Y.A robust electricity price forecasting framework based on heteroscedastic temporal Convolutional Network2024Scopus
H. -Y. Su; G. -Z. LiaoA Spike-Resilient Temporal-Adaptive Neural Framework for Day-Ahead Electricity Price Interval Forecasting2025IEEE
Lee, Chang-Min; Kwang, Song, Sung; Sungwook, Chung,A Study on the Prediction of Cabbage Price Using Ensemble Voting Techniques2022WOS
Li, J.; Wen, X.; Jia, L.; Cao, R.; Zhang, X.; Hao, Y.; Gao, C.; Cao, R.A Transformer-based model fusing temporal dependence and variable correlation for short and medium-term electricity price forecasting2025Scopus
M. T. Luc; R. Senkerik; T. K. DangAdvanced Cryptocurrency Price Prediction: Incorporating Time-Reversed ARIMA and BVIC2025IEEE
F. Chauvet; L. Bellatreche; C. A. Santos SilvaAI Approaches for Electricity Price Forecasting in Stable/Unstable Markets: EU Improvement Project2022IEEE
K. Lakshmi.; P. V. Narayana; P. D. Bhavani; V. V. S. Madhavacharyulu; N. Lavanya; S. PousiaAn Enhanced Regression Technique for House Price Prediction2022IEEE
Heistrene, L.; Belikov, J.; Baimel, D.; Katzir, L.; Machlev, R.; Levy, K.; Mannor, S.; Levron, Y.An Improved and Explainable Electricity Price Forecasting Model via SHAP-Based Error Compensation Approach2025Scopus
Cheng, X; Wu, P; Liao, SS; Wang, XLAn integrated model for crude oil forecasting: Causality assessment and technical efficiency2023WoS
Öner, G.; Goksu, B.; Acikgoz, E.Anomaly levels of factors affecting the estimated price of a second-hand ship2025Scopus
Y. Huang; L. Huang; Y. Zhu; G. Lei; A. Wang; J. ZhuApplication of Large Language Models in Intelligent Preprocessing and Forecasting of Electricity Price2025IEEE
Stancu, S.; Isaic-Maniu, A.; Bodea, C.-N.; Muscalu, M.S.; Bala, D.E.Application of Machine Learning Techniques in Natural Gas Price Modeling. Analyses, Comparisons, and Predictions for Romania2024Scopus
Atsalaki, I; Atsalakis, GS; Melas, KD; Michail, NABaltic dry index forecasting using a neuro-fuzzy inference system2025WoS
Singh, A.; Shukla, B.; Jos, J.Comparative Analysis of CPI prediction for India using Statistical methods and Neural Networks2023Scopus
Arshad, J.; Akter, A.; Akter, T.; Ghosh, K.P.; Singha, A.Cryptocurrency Price Prediction and Security Challenges in Machine Learning Approach2026Scopus
Akter, K.; Rahman, M.A.; Islam, M.R.; Tabassum, F.Data reconstruction-enabled robust hybrid models for intraday electricity price forecasting against data contamination attacks2026Scopus
Liu, SY; Jiang, YC; Lin, ZZ; Wen, FS; Ding, Y; Yang, LData-driven Two-step Day-ahead Electricity Price Forecasting Considering Price Spikes2023WoS
Tan, Y.Q.; Shen, Y.X.; Yu, X.Y.; Lu, X.Day-ahead electricity price forecasting employing a novel hybrid frame of deep learning methods: A case study in NSW, Australia2023Scopus
M. Miletic; I. Pavic; H. Pandzic; T. CapuderDay-ahead Electricity Price Forecasting Using LSTM Networks2022IEEE
M. N. Ahsan; M. Moniruzzaman; S. JahanDeep Learning-Based Stock Price Forecasting and Anomaly Detection: A Data-Driven Approach2025IEEE
B. H. Impa; V. Darshan; B. Soradi; D. Tejas; D. Prakash; M. V. MadhusudhanDevelopment of AI-ML-Based Models for Predicting Prices of Agri-Horticultural Commodities2025IEEE
Zhao, Y.; Feng, C.; Xu, N.; Peng, S.; Liu, C.Early warning of exchange rate risk based on structural shocks in international oil prices using the LSTM neural network model2023Scopus
Nafea, AA; Alawi, OA; Doost, ZH; Alani, MM; Ahmadullah, AB; Deo, RC; Yaseen, ZMElectricity load and price forecasting in Spain: A hybrid deep learning framework leveraging temporal and seasonal dynamics2026WoS
S. Albahli; M. Shiraz; N. AyubElectricity Price Forecasting for Cloud Computing Using an Enhanced Machine Learning Model2020IEEE
L. Han; C. Ban; C. Zhang; L. LiuElectricity Price Forecasting in Power Markets Based on Machine Learning2024IEEE
A. K. Prajapati; S. K. Srivastava; A. NarainElectricity Price Forecasting: A Bibliographical Review2020IEEE
L. Gong; J. Fu; J. Sun; H. Tang; S. Gu; F. WangElectricity Price Prediction Based on QPSO-CNN-LSTM in the Spot Market Environment2024IEEE
Sage, M.; Campbell, J.; Zhao, Y.F.ENHANCING BATTERY STORAGE ENERGY ARBITRAGE WITH DEEP REINFORCEMENT LEARNING AND TIME-SERIES FORECASTING2024Scopus
Cerasa, A.; Zani, A.Enhancing electricity price forecasting accuracy: A novel filtering strategy for improved out-of-sample predictions2025Scopus
Mehra, I.; Tyagi, M.K.; Kanthalia, S.; Shree, T.; Chaturvedi, R.K.; Singh, P.Evaluating LSTM and GRU Models for Cryptocurrency Price Forecasting in Financial Markets2023Scopus
Sogukpinar, Fatih; Erkal, Gokhan; Ozer, HuseyinEvaluation of renewable energy policies in Turkey with sectoral electricity demand forecasting2023WOS
N. Manjunathan; U. S. S; R. Nithyanandhan; S. Swetha; M. B; N. MExploring Machine Learning and Advanced Modeling Techniques in Financial Data Science2025IEEE
Yahia, Achraf; Mouhssine, Yassine; El Alaoui, Abdelkader; El Alaoui, Said OuatikExploring machine learning-based methods for anomalies detection: evidence from cryptocurrencies2025WOS
D. Kumar; A. Kumar; R. K. Burman; P. Kumar; A. Sinha; A. K. MahatoFair Price Prediction of Vegetables for Farmers Using Machine Learning2025IEEE
Asl, MG; Adekoya, OB; Rashidi, MM; Doudkanlou, MG; Dolatabadi, AForecast of Bayesian-based dynamic connectedness between oil market and Islamic stock indices of Islamic oil-exporting countries: Application of the cascade-forward backpropagation network2022WoS
A. Sridhar; M. Karhunen; S. Honkapuro; F. RuizForecast or Nowcast to Predict Electricity Prices? The Role of Open Data2024IEEE
Akhter, T.; Ratna, T.S.; Ahmed, F.; Babu, M.A.; Hossain, S.F.A.Forecasting and unveiling the impeded factors of total export of Bangladesh using nonlinear autoregressive distributed lag and machine learning algorithms2024Scopus
Parvini, N.; Abdollahi, M.; Seifollahi, S.; Ahmadian, D.Forecasting Bitcoin returns with long short-term memory networks and wavelet decomposition: A comparison of several market determinants2022Scopus
Swei, OmarForecasting Infidelity: Why Current Methods for Predicting Costs Miss the Mark2020WOS
Stathakis, E.; Papadimitriou, T.; Gogas, P.Forecasting price spikes in electricity markets2021Scopus
Liu, L.; Bai, F.; Su, C.; Ma, C.; Yan, R.; Li, H.; Sun, Q.; Wennersten, R.Forecasting the occurrence of extreme electricity prices using a multivariate logistic regression model2022Scopus
Hirota, S.; Cui, J.; Fang, X.; Oozeki, T.; Ueda, Y.Forecasting Wholesale Electricity Market Prices Considering Bidding Conditions Using Price Sensitivity2025Scopus
Ren, Xiaohang; Xu, Weixi; Duan, KunFourier transform-based LSTM stock prediction model under oil shocks2022WOS
Chakraborty, S.; Jagabathula, S.; Subramanian, L.; Venkataraman, A.Frontiers in Operations: News Event-Driven Forecasting of Commodity Prices2024Scopus
Nguyen, Hoach The; Al-Sumaiti, Ameena Saad; Turitsyn, Konstantin; Li, Qifeng; El Moursi, Mohamed ShawkyFurther Optimized Scheduling of Micro Grids via Dispatching Virtual Electricity Storage Offered by Deferrable Power-Driven Demands2020WOS
Hegde, G.; Hulipalled, V.R.; Simha, J.B.Fuzzified input data tuning for agriculture commodities price prediction2024Scopus
Shahzad, Umer; Mohammed, Kamel Si; Schneider, Nicolas; Faggioni, Francesca; Papa, ArmandoGDP responses to supply chain disruptions in a post-pandemic era: Combination of DL and ANN outputs based on Google Trends2023WOS
Yang, H.; Schell, K.R.GHTnet: Tri-Branch deep learning network for real-time electricity price forecasting2022Scopus
H. Yang; K. R. SchellHFNet: Forecasting Real-Time Electricity Price via Novel GRU Architectures2020IEEE
Vahedi, N.; Jolfaei, A.; Ur Rehman, S.; Ravi, R.S.Hybrid Deep Learning Model for Electricity Price Forecasting in Renewable-Dominated Markets: The Case of South Australia2025Scopus
I. U. Khan; M. JamilHybrid Long Short-Term Memory (LSTM) and Exponentially Weighted Moving Average (EWMA) Model for Accurate and Scalable Electricity Price Forecasting2025IEEE
C. Mylordick; L. A. WulandhariImproving Stock Price Prediction by Utilizing Anomaly Detection and Handling for Top 5 FMCG Company in Indonesia2025IEEE
Deng, Z.; Liu, C.; Zhu, Z.Inter-hours rolling scheduling of behind-the-meter storage operating systems using electricity price forecasting based on deep convolutional neural network2021Scopus
Wang, J; Wang, XY; Wang, XInternational oil shocks and the volatility forecasting of Chinese stock market based on machine learning combination models2024WoS
D. D. Prastyo; I. Yuniarti; Setiawan; S. P. RahayuInterval Forecasting Using GEV Regression With Stochastic Search Variable Selection: A Simulation Study and Its Application to Forecast Imported Meat Price Range2025IEEE
Sinclair, M.; Shepley, A.J.; Hajati, F.Learning the Grid: Transformer Architectures for Electricity Price Forecasting in the Australian National Market2026Scopus
Irakora Kalisimbi, E.H.; Zhou, Y.; Uwimana, E.; Sall, N.M.Leveraging Deep Learning and Hybrid Models for Enhanced Electricity Load and Price Forecasting in Deregulated Utilities Market2025Scopus
A. Javed; R. R. DerakhshaniMachine Learning Ensembles for Grid Congestion Price Forecasting2022IEEE
Candila, Vincenzo; Petrella, Lea; Andreani, MilaMixed-frequency Quantile Regression Forests for Value-at-Risk forecasting2025WOS
Wang, ZY; Mae, M; Yamane, T; Ajisaka, M; Nakata, T; Matsuhashi, RNovel Custom Loss Functions and Metrics for Reinforced Forecasting of High and Low Day-Ahead Electricity Prices Using Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) and Ensemble Learning2024WoS
Kim, Y.; Ghosh, A.; Topal, E.; Chang, P.Performance of different models in iron ore price prediction during the time of commodity price spike2023Scopus
C. Yan; H. RuanPower Load and Price Forecasting Based on Spatiotemporal Decoupling and Explainable Artificial Intelligence Interaction Perception2025IEEE
Zou, YZ; Herremans, DPreBit-A multimodal model with Twitter FinBERT embeddings for extreme price movement prediction of Bitcoin2023WoS
S. VolkovaPrecious Metals Price Forecasting Using Neural Networks2024IEEE
Pujitha, M.V.; Kumar, K.S.; Sree, D.I.; Devi, P.Y.Predicting Bitcoin Prices for the Next 30 Days Using LSTM-Based Time Series Analysis2023Scopus
J. Schmitz; P. RanganathanPredicting Locational-Based Marginal Pricing (LBMP) Using Ridge Regression2024IEEE
Al-Sulaiman, T.Predicting reactions to anomalies in stock movements using a feed-forward deep learning network2022Scopus
Yang, Yu; Wang, Jun; Wang, BinPrediction model of energy market by long short-term memory with random system and complexity evaluation2020WOS
Rantonen, M.; Korpihalkola, J.Prediction of Spot Prices in Nord Pool’s Day-Ahead Market Using Machine Learning and Deep Learning2020Scopus
Ehsani, B.; Pineau, P.-O.; Charlin, L.Price forecasting in the Ontario electricity market via TriConvGRU hybrid model: Univariate vs. multivariate frameworks2024Scopus
C. Zhang; L. Gong; Y. Peng; Y. FuPrice Spike Classification and Regression Using A Hybrid Oversampling Method2022IEEE
Alruqimi, M.; Di Persio, L.Probabilistic Forecasting of Crude Oil Prices Using Conditional Generative Adversarial Network Model with Lévy Process2025Scopus
Narajewski, M.Probabilistic Forecasting of German Electricity Imbalance Prices2022Scopus
Negri, P.; Ramos, P.; Breitkopf, M.Regional Commodities Price Volatility Assessment Using Self-driven Recurrent Networks2021Scopus
N. Wilkins; M. Johnson; I. NwoguRegression with Uncertainty Quantification in Large Scale Complex Data2022IEEE
Wang Yunfang; Shi Yi; Chen LihuaResearch on Forecast Model of Fresh Fruit and Vegetable Logistics Demand Based on BA-SVR Hybrid Model2024WOS
X. Chen; S. ZhaoResearch on Noise Suppression and Signal Enhancement in Electricity Price Anomaly Detection2025IEEE
Sun, Yunpeng; Guo, Jin; Shan, Shan; Khan, Yousaf AliRETRACTED: Wheat Futures Prices Prediction in China: A Hybrid Approach (Retracted Article)2021WOS
Yi, S.; Chi, J.; Shi, Y.; Zhang, C.Robust stock trend prediction via volatility detection and hierarchical multi-relational hypergraph attention2025Scopus
Y. Deng; K. Zheng; X. Feng; Q. ChenScenario Analysis and Anomaly Detection of Locational Marginal Price in Electricity Market2020IEEE
Deng, S.; Inekwe, J.; Smirnov, V.; Wait, A.; Wang, C.Seasonality in deep learning forecasts of electricity imbalance prices2024Scopus
Prajesh, A.; Jain, P.; Dhingra, S.; Rishabh; Malik, A.; Agrawal, S.Short Term Electricity Price Forecasting by Optimized LSTM Model of Deep Learning with Genetic Algorithm2023Scopus
Makri, E.; Koskinas, I.; Tsolakis, A.C.; Ioannidis, D.; Tzovaras, D.Short Term Net Imbalance Volume Forecasting Through Machine and Deep Learning: A UK Case Study2021Scopus
Zhang, C.; Fu, Y.; Gong, L.Short-Term Electricity Price Forecast Using Frequency Analysis and Price Spikes Oversampling2023Scopus
I. Shah; S. Akbar; T. Saba; S. Ali; A. RehmanShort-Term Forecasting for the Electricity Spot Prices With Extreme Values Treatment2021IEEE
Perla, S.; Bisoi, R.; Dash, P.K.; Rout, A.K.Short-term forecasting of electricity price using ensemble deep kernel-based random vector functional link network2025Scopus
Mari, C.; Mari, E.Stochastic DNN-based models meet hidden Markov models: a challenge on natural gas prices at the Henry Hub2025Scopus
Suliman, M.S.; Farzaneh, H.Synthesizing the market clearing mechanism based on the national power grid using hybrid of deep learning and econometric models: Evidence from the Japan Electric Power Exchange (JEPX) market2023Scopus
N. Lee; P. MandalTemporal Convolutional Network Ensemble for Day-Ahead and Week-Ahead Electricity Price Forecasting2023IEEE
Tissaoui, Kais; Zaghdoudi, Taha; Boubaker, Sahbi; Hkiri, Besma; Talbi, MariemTesting the Nonlinear Long- and Short-Run Distributional Asymmetries Effects of Bitcoin Prices on Bitcoin Energy Consumption: New Insights through the QNARDL Model and XGBoost Machine-Learning Tool2024WOS
Abdou, Hussein A.; Elamer, Ahmed A.; Abedin, Mohammad Zoynul; Ibrahim, Bassam A.The impact of oil and global markets on Saudi stock market predictability: A machine learning approach2024WOS
Wang, JZ; Niu, XS; Zhang, LF; Liu, ZK; Wei, DXThe influence of international oil prices on the exchange rates of oil exporting countries: Based on the hybrid copula function2022WoS
Li, Yuze; Jiang, Shangrong; Li, Xuerong; Wang, ShouyangThe role of news sentiment in oil futures returns and volatility forecasting: Data-decomposition-based deep learning approach2021WOS
Kmytiuk, T.; Majore, G.; Bilyk, T.TIME SERIES FORECASTING OF PRICE OF THE AGRICULTURAL PRODUCTS USING DATA SCIENCE2024Scopus
M. S. Shivalini; P. K. Gandra; S. Alamuri; S. S. Gupta; P. STransforming Financial Market Prediction Through Generative AI-Powered Data Management2024IEEE
Asirim, Ö.E.Trend Identification using LSTM vs. Convolutional Neural Networks on the NASDAQ market2025Scopus
S. Hiremath; S. Gowthami; N. Jahnavi; K. R. B. P; K. KumariUnfair Trading Detection System-A Hybrid Approach2025IEEE
Puka, Radoslaw; Lamasz, BartoszUsing Artificial Neural Networks to Find Buy Signals for WTI Crude Oil Call Options2020WOS
Zafar, H.; Kapetanakis, S.Visualizing Bitcoin’s Evolution: Dimensionality Reduction and Stylized Facts Techniques2025Scopus
D. Thombare; A. Khadake; R. Kumbhar; A. Ingle; M. S. Deshmukh; M. WadhwaniWeb Scraping-Based Cryptocurrency Prediction and Analysis2025IEEE
Sridharan, V.; Tuo, M.; Li, X.Wholesale Electricity Price Forecasting Using Integrated Long-Term Recurrent Convolutional Network Model2022Scopus

References

  1. Higgs, H. Modelling price and volatility inter-relationships in the Australian wholesale spot electricity markets. Energy Econ. 2009, 31, 748–756. [Google Scholar] [CrossRef] [Scilit]
  2. Bao, T.; Mahdavi, N.; McCarthy, C.; Rezazadegan, D. Electricity price spike forecasting via multivariate adaptive visibility graph neural networks. Electr. Power Syst. Res. 2026, 255, 112695. [Google Scholar] [CrossRef] [Scilit]
  3. Jaimes, D.M.; López, M.Z.; Zareipour, H.; Quashie, M. A Hybrid Model for Multi-Day-Ahead Electricity Price Forecasting considering Price Spikes. Forecasting 2023, 5, 499–521. [Google Scholar] [CrossRef] [Scilit]
  4. Paskaleva, M.; Vasenska, I. The BTC Price Prediction Paradox Through Methodological Pluralism. Risks 2025, 13, 195. [Google Scholar] [CrossRef] [Scilit]
  5. Peng, Y.; Wang, Z.; Castillo, I.; Gunnell, L.; Jiang, S. A New Modeling Framework for Real-Time Extreme Electricity Price Forecasting. IFAC-Pap. Conf. Pap. 2024, 58, 899–904. [Google Scholar] [CrossRef] [Scilit]
  6. Jiang, Z.; Wang, J.; Zhang, T.; Li, G.; Zhou, M. Deep Learning-Based Hybrid Model for Forecasting Locational Marginal Prices; Institute of Electrical and Electronics Engineers Inc.: Piscataway NY, USA, 2020; pp. 1733–1738. [Google Scholar] [CrossRef] [Scilit]
  7. Dinler, A. A Review of Balancing Price Forecasting in the Context of Renewable-Rich Power Systems, Highlighting Profit-Aware and Spike-Resilient Approaches. Energies 2025, 18, 6460. [Google Scholar] [CrossRef] [Scilit]
  8. Jasiński, T. A Review of Recent Trends in Electricity Price Forecasting Using Deep Learning Techniques. Energies 2025, 18, 6422. [Google Scholar] [CrossRef] [Scilit]
  9. Jiang, L.; Hu, G. A Review on Short-Term Electricity Price Forecasting Techniques for Energy Markets; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NY, USA, 2018; pp. 937–944. [Google Scholar] [CrossRef] [Scilit]
  10. Lago, J.; Marcjasz, G.; De Schutter, B.; Weron, R. Forecasting day-ahead electricity prices: A review of state-of-the-art algorithms, best practices and an open-access benchmark. Appl. Energy 2021, 293, 116983. [Google Scholar] [CrossRef] [Scilit]
  11. O’connor, C.; Bahloul, M.; Prestwich, S.; Visentin, A. A review of electricity price forecasting models in the day-ahead, intra-day, and balancing markets. Energies 2025, 18, 3097. [Google Scholar] [CrossRef] [Scilit]
  12. Rashid, N.A.; Ismail, M.T. A Review: Predictive Models and Behaviour of Cryptocurrencies Price. J. Adv. Res. Appl. Sci. Eng. Technol. 2024, 48, 148–167. [Google Scholar] [CrossRef] [Scilit]
  13. Wen, Q.; Zhou, T.; Zhang, C.; Chen, W.; Ma, Z.; Yan, J.; Sun, L. Transformers in time series: A survey. arXiv 2022, arXiv:2202.07125. [Google Scholar]
  14. Shyalika, C.; Wickramarachchi, R.; Sheth, A.P. A comprehensive survey on rare event prediction. ACM Comput. Surv. 2024, 57, 1–39. [Google Scholar] [CrossRef] [Scilit]
  15. Carmo, J.E.D.; Zareipour, H. Leveraging Temporal Patterns in Electricity Price Forecasting with Cluster-Based Feature Engineering and Temporal Fusion Transformers. In Proceedings of the 2024 IEEE 42nd Central America and Panama Convention (CONCAPAN XLII), San Jose, Costa Rica, 27–29 November 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
  16. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, 71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Lim, B.; Arık, S.Ö.; Loeff, N.; Pfister, T. Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 2021, 37, 1748–1764. [Google Scholar] [CrossRef] [Scilit]
  18. Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. Proc. AAAI Conf. Artif. Intell. 2021, 35, 11106–11115. [Google Scholar] [CrossRef] [Scilit]
  19. Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Adv. Neural Inf. Process. Syst. 2021, 34, 22419–22430. [Google Scholar] [CrossRef] [Scilit]
  20. Karim, E.; Ahmed, S. A Deep Learning-Based Approach for Stock Price Prediction Using Bidirectional Gated Recurrent Unit and Bidirectional Long Short Term Memory Model; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NY, USA, 2020. [Google Scholar] [CrossRef] [Scilit]
  21. Yang, W.; Wang, R.; Wang, B. Detection of Anomaly Stock Price Based on Time Series Deep Learning Models; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NY, USA, 2020; pp. 110–114. [Google Scholar] [CrossRef] [Scilit]
  22. Neyret, M.; Ouaggag, J.; Allain, C. Trading Desk Behavior Modeling via LSTM for Rogue Trading Fraud Detection; Hammoudi, S., Quix, C., Bernardino, J., Eds.; SciTePress: Setúbal, Portugal, 2020; pp. 143–150. [Google Scholar]
  23. Yeung, J.F.K.A.; Wei, Z.-K.; Chan, K.Y.; Lau, H.Y.K.; Yiu, K.-F.C. Jump detection in financial time series using machine learning algorithms. Soft Comput. 2020, 24, 1789–1801. [Google Scholar] [CrossRef] [Scilit]
  24. Jakaša, T.; Andročec, I.; Sprčić, P. Electricity Price Forecasting—ARIMA Model Approach; IEEE: Piscataway, NJ, USA, 2011; pp. 222–225. [Google Scholar] [CrossRef] [Scilit]
  25. Voronin, S.; Partanen, J.; Kauranne, T. A hybrid electricity price forecasting model for the Nordic electricity spot market. Int. Trans. Elecr. Energy Sys. 2014, 24, 736–760. [Google Scholar] [CrossRef] [Scilit]
  26. Lux, M.; Härdle, W.K.; Lessmann, S. Data driven value-at-risk forecasting using a SVR-GARCH-KDE hybrid. Comput. Stat. 2020, 35, 947–981. [Google Scholar] [CrossRef] [Scilit]
  27. Razak, I.A.B.W.A.; Abidin, I.b.Z.; Siah, Y.K.; Rahman, T.K.B.A.; Lada, M.Y.; Ramani, A.N.B.; Nasir, M.N.M.; Ahmad, A.B. Support vector machine for day ahead electricity price forecasting. In Proceedings of the AIP Conference Proceedings, Penang, Malaysia, 28–30 May 2014; Ramli, M.F., Roslan, N., Junoh, A.K., Masnan, M.J., Kharuddin, M.H., Eds.; American Institute of Physics Inc.: College Park, MD, USA, 2015; Volume 1660. [Google Scholar] [CrossRef] [Scilit]
  28. Koban, V.; Zlatar, I.; Pantoš, M.; Omladič, M. A remark on forecasting spikes in electricity prices. In Proceedings of the International Conference on the European Energy Market (EEM), Lisbon, Portugal, 19–22 May 2015; IEEE Computer Society: Washington, DC, USA, 2015; Volume 2015. [Google Scholar] [CrossRef] [Scilit]
  29. Shrivastava, N.A.; Panigrahi, B.K.; Lim, M.H. Electricity price classification using extreme learning machines. Neural Comput. Appl. 2016, 27, 9–18. [Google Scholar] [CrossRef] [Scilit]
  30. Stathakis, E.; Papadimitriou, T.; Gogas, P. Forecasting price spikes in electricity markets. Rev. Econimic Anal. 2020, 12, 1–24. [Google Scholar] [CrossRef] [Scilit]
  31. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar] [CrossRef] [Scilit]
  32. Miletic, M.; Pavic, I.; Pandzic, H.; Capuder, T. Day-ahead Electricity Price Forecasting Using LSTM Networks. In Proceedings of the 2022 7th International Conference on Smart and Sustainable Technologies (SpliTech), Split/Bol, Croatia, 5–8 July 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  33. Gong, L.; Fu, J.; Sun, J.; Tang, H.; Gu, S.; Wang, F. Electricity Price Prediction Based on QPSO-CNN-LSTM in the Spot Market Environment. In Proceedings of the 2024 5th International Conference on Information Science, Parallel and Distributed Systems (ISPDS), Guangzhou, China, 31 May–2 June 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 343–346. [Google Scholar] [CrossRef] [Scilit]
  34. Nafea, A.A.; Alawi, O.A.; Doost, Z.H.; AlAni, M.; Ahmadullah, A.B.; Deo, R.C.; Yaseen, Z.M. Electricity load and price forecasting in Spain: A hybrid deep learning framework leveraging temporal and seasonal dynamics. EXPERT Syst. Appl. 2025, 299, 130123. [Google Scholar] [CrossRef] [Scilit]
  35. Perla, S.; Bisoi, R.; Dash, P.K.; Rout, A.K. Short-term forecasting of electricity price using ensemble deep kernel based random vector functional link network. Appl. Soft Comput. 2025, 174, 113012. [Google Scholar] [CrossRef] [Scilit]
  36. Su, H.-Y.; Liao, G.-Z. A Spike-Resilient Temporal-Adaptive Neural Framework for Day-Ahead Electricity Price Interval Forecasting. IEEE Trans. Power Syst. 2025, 40, 4411–4414. [Google Scholar] [CrossRef] [Scilit]
  37. Shi, W.; Wang, Y.F. A robust electricity price forecasting framework based on heteroscedastic temporal Convolutional Network. Int. J. Electr. Power Energy Syst. 2024, 161, 110177. [Google Scholar] [CrossRef] [Scilit]
  38. Yang, H.; Schell, K.R. GHTnet: Tri-Branch deep learning network for real-time electricity price forecasting. Energy 2022, 238, 122052. [Google Scholar] [CrossRef] [Scilit]
  39. Kırat, O.; Cicek, A.; Yerlikaya, T. A new artificial intelligence-based system for optimal electricity arbitrage of a second-life battery station in day-ahead markets. Appl. Sci. 2024, 14, 10032. [Google Scholar] [CrossRef] [Scilit]
  40. Sage, M.; Campbell, J.; Zhao, Y.F. Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting; American Society of Mechanical Engineers (ASME): New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
  41. Negri, P.; Ramos, P.; Breitkopf, M. Regional Commodities Price Volatility Assessment Using Self-driven Recurrent Networks. In Lecture Notes in Computer Science (LNCS); Tavares, J.M., Papa, J.P., González Hidalgo, M., Eds.; Springer Science and Business Media Deutschland GmbH: Berlin/Heidelberg, Germany, 2021; Volume 12702, pp. 361–370. [Google Scholar] [CrossRef] [Scilit]
  42. Yi, S.; Chi, J.; Shi, Y.; Zhang, C. Robust stock trend prediction via volatility detection and hierarchical multi-relational hypergraph attention. Knowl.-Based Syst. 2025, 329, 114283. [Google Scholar] [CrossRef] [Scilit]
  43. Deng, Y.; Zheng, K.; Feng, X.; Chen, Q. Scenario Analysis and Anomaly Detection of Locational Marginal Price in Electricity Market; IEEE: Piscataway, NJ, USA, 2020; Volume 2020, pp. 2378–2384. [Google Scholar] [CrossRef] [Scilit]
  44. Alruqimi, M.; Di Persio, L. Probabilistic Forecasting of Crude Oil Prices Using Conditional Generative Adversarial Network Model with Lévy Process. Mathematics 2025, 13, 307. [Google Scholar] [CrossRef] [Scilit]
  45. Narajewski, M. Probabilistic Forecasting of German Electricity Imbalance Prices. Energies 2022, 15, 4976. [Google Scholar] [CrossRef] [Scilit]
  46. Shi, W.; Wang, Y.; Chen, Y.; Ma, J. An effective Two-Stage Electricity Price forecasting scheme. Electr. Power Syst. Res. 2021, 199, 107416. [Google Scholar] [CrossRef] [Scilit]
  47. Huang, X.; Liu, G.; Huan, J.; Luo, S.; Qiu, J.; Qin, F.; Xu, Y. Combined use of long short-term memory neural network and quantum computation for hierarchical forecasting of locational marginal prices. Energy Convers. Econ. 2025, 6, 26–38. [Google Scholar] [CrossRef] [Scilit]
  48. Wu, J.; Wu, L.; Xu, Z.; Qiao, X.; Guan, X. Dynamic pricing and prices spike detection for industrial park with coupled electricity and thermal demand. IEEE Trans. Autom. Sci. Eng. 2022, 19, 1326–1337. [Google Scholar] [CrossRef] [Scilit]
  49. Shahinzadeh, H.; Nafisi, H.; Bagheri, M.; Mehrabani-Najafabadi, S.; Hejazi, M.A.; Jurado, F. Forecasting Electricity Price Spikes in Competitive Markets Using a Hybrid Deep Learning Framework with SHAP and Grey Wolf Optimization. In Proceedings of the 2025 6th International Conference on Optimizing Electrical Energy Consumption (OEEC), Najafabad, Iran, Islamic, 25–26 February 2025; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2025; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  50. Farhoumandi, M.; Bahramirad, S.; Alabdulwahab, A.; Shahidehpour, M.; Rahimi, F.; Ipakchi, A.; Albuyeh, F.; Mokhtari, S. Fusion of deep learning and machine learning methods for hourly locational marginal price forecast in power systems. iEnergy 2025, 4, 193–204. [Google Scholar] [CrossRef] [Scilit]
  51. Mao, R.; Zheng, Z.; Lai, S.; Tao, Y.; Dong, Z.Y.; Li, T.; Huang, Z.; Qiu, J. Large Language Model Based Data Augmentation for Peak Electricity Price Forecasting and Battery Energy Storage Arbitrage. In Proceedings of the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria, 5–8 October 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 7762–7767. [Google Scholar] [CrossRef] [Scilit]
  52. Liu, C.; Cai, L.; Dalzell, G.; Mills, N. Large Language Model for Extreme Electricity Price Forecasting in the Australia Electricity Market. In Proceedings of the IECON 2024—50th Annual Conference of the IEEE Industrial Electronics Society, Chicago, IL, USA, 3–6 November 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  53. Lee, N.; Ngu, N.; Sahdev, H.S.; Motaganahall, P.; Chowdhury, A.M.S.; Xi, B.; Shakarian, P. Metal Price Spike Prediction via a Neurosymbolic Ensemble Approach. In Proceedings of the 2024 IEEE International Conference on Data Mining Workshops (ICDMW), Abu Dhabi, United Arab Emirates, 9 December 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 106–110. [Google Scholar] [CrossRef] [Scilit]
  54. Li, Y.; Li, C.; Chen, G.; Zhou, X.; Dong, Z. Multi-Task Graph Adaptive Learning for Multivariate Electricity Price Short-Term Forecasting in Australia’s National Electricity Market. IEEE Trans. Power Syst. 2025, 40, 530–542. [Google Scholar] [CrossRef] [Scilit]
  55. Wu, Y.; Wang, S.; Fu, X. mWDN-Transformer: Utilizing Cyclical Patterns for Short-Term Forecasting of Limit Order Book in China Markets. In Proceedings of the 2024 5th International Conference on Computing, Networks and Internet of Things, Tokyo Japan, 24–26 May 2024; Association for Computing Machinery: New York, NY, USA, 2024; pp. 175–181. [Google Scholar] [CrossRef] [Scilit]
  56. Alvarez, A.; Luo, W.; Fryer, S. Ultra-short term wholesale electricity price forecasting through deep learning. In Proceedings of the 2021 IEEE PES Innovative Smart Grid Technologies—Asia (ISGT Asia), Brisbane, Australia, 5–8 December 2021; IEEE: Piscataway, NY, USA, 2021. [Google Scholar] [CrossRef] [Scilit]
  57. Kochliaridis, V.; Papadopoulou, A.; Vlahavas, I. UNSURE—A machine learning approach to cryptocurrency trading. Appl. Intell. 2024, 54, 5688–5710. [Google Scholar] [CrossRef] [Scilit]
  58. Zamudio Lopez, M. Modeling and Forecasting Extreme Electricity Prices in Competitive Markets. 2026. Available online: https://ucalgary.scholaris.ca/items/c6a628b1-7342-4e27-a42e-d2f9cb871d90 (accessed on 26 January 2026).
  59. Masum, K.M.A.; M.Azad, A.K.; Debi, C.R.; Birǎu, R.; Popescu, V.; Ionascu, C.M. From Topology to Geometry: A Neural Ricci Flow Framework for Predicting Flash Crashes and Contagion. IEEE Access 2026, 14, 26912–26934. [Google Scholar] [CrossRef] [Scilit]
  60. Liu, S.; Jiang, Y.; Lin, Z.; Wen, F.; Ding, Y.; Yang, L. Data-driven Two-step Day-ahead Electricity Price Forecasting Considering Price Spikes. J. Mod. Power Syst. Clean Energy 2023, 11, 523–533. [Google Scholar] [CrossRef] [Scilit]
  61. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Sinclair, M.; Shepley, A.J.; Hajati, F. Learning the Grid: Transformer Architectures for Electricity Price Forecasting in the Australian National Market. Appl. Sci. 2025, 16, 75. [Google Scholar] [CrossRef] [Scilit]
  63. Khan, I.U.; Jamil, M. Hybrid Long Short-Term Memory (LSTM) and Exponentially Weighted Moving Average (EWMA) Model for Accurate and Scalable Electricity Price Forecasting. In Proceedings of the 2025 IEEE 13th International Conference on Smart Energy Grid Engineering (SEGE), Oshawa, ON, Canada, 18–20 August 2025; IEEE: Piscataway, NY, USA, 2025; pp. 107–111. [Google Scholar] [CrossRef] [Scilit]
  64. Zhou, Z.; Basker, R.; Yeung, D.-Y. Graph neural networks for multivariate time-series forecasting via learning hierarchical spatiotemporal dependencies. Eng. Appl. Artif. Intell. 2025, 147, 110304. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Conceptual positioning of the review within the broader forecasting landscape. Rare event prediction is a specialised subset of market forecasting, further constrained to time-series modelling, with this review focusing specifically on deep learning-based approaches.
Figure 1. Conceptual positioning of the review within the broader forecasting landscape. Rare event prediction is a specialised subset of market forecasting, further constrained to time-series modelling, with this review focusing specifically on deep learning-based approaches.
Forecasting 08 00052 g001
Figure 3. Distribution of included studies by market domain. Number shown represent the count of studies.
Figure 3. Distribution of included studies by market domain. Number shown represent the count of studies.
Forecasting 08 00052 g003
Figure 4. A stacked LSTM where the input sequence is first processed by a lower LSTM layer, whose outputs at each time step are passed to a second LSTM layer. The upper layer produces predictions at each time step, allowing the model to learn hierarchical temporal representations. At time step t , the LSTM takes input x t along with the previous hidden state h t 1 and cell state C t 1 , and produces updated states h t and C t . The sigmoid ( σ ) and tanh functions control three gates that regulate information flow: deciding what to forget, what new information to add, and what to output. Circles with × and + denote element-wise multiplication and addition, respectively.
Figure 4. A stacked LSTM where the input sequence is first processed by a lower LSTM layer, whose outputs at each time step are passed to a second LSTM layer. The upper layer produces predictions at each time step, allowing the model to learn hierarchical temporal representations. At time step t , the LSTM takes input x t along with the previous hidden state h t 1 and cell state C t 1 , and produces updated states h t and C t . The sigmoid ( σ ) and tanh functions control three gates that regulate information flow: deciding what to forget, what new information to add, and what to output. Circles with × and + denote element-wise multiplication and addition, respectively.
Forecasting 08 00052 g004
Figure 5. Graph neural network architecture for multivariate time-series forecasting. Colours indicate major processing stages: CNN feature extraction (purple), data transformation operations (blue), and graph convolution processing (green). Circles denote graph nodes and squares denote node features.
Figure 5. Graph neural network architecture for multivariate time-series forecasting. Colours indicate major processing stages: CNN feature extraction (purple), data transformation operations (blue), and graph convolution processing (green). Circles denote graph nodes and squares denote node features.
Forecasting 08 00052 g005
Figure 6. Encoder–decoder transformer architecture for electricity price forecasting. Historical observations are processed through stacked self-attention encoder (pink) layers, while the decoder (green) combines forecast horizons and contextual inputs to generate future price predictions.
Figure 6. Encoder–decoder transformer architecture for electricity price forecasting. Historical observations are processed through stacked self-attention encoder (pink) layers, while the decoder (green) combines forecast horizons and contextual inputs to generate future price predictions.
Forecasting 08 00052 g006
Table 1. Comparison of existing review studies relevant to forecasting, rare-event prediction, and spike forecasting.
Table 1. Comparison of existing review studies relevant to forecasting, rare-event prediction, and spike forecasting.
ReviewDomainDeep Learning FocusExplicit Spike FocusRare-Event FocusCross-Domain
[9]ElectricityLimitedNoNoNo
[10]ElectricityPartialNoNoNo
[7]ElectricityPartialYesPartialNo
[8]ElectricityYesNoNoNo
[11]ElectricityYesLimitedNoNo
[12]CryptocurrencyLimitedNoNoNo
[13]Time-seriesYesNoNoYes
[14]Rare-event predictionNoNoYesYes
Present ReviewMulti-domain spike forecastingYesYesYesYes
Table 2. Search terms used to obtain articles for review.
Table 2. Search terms used to obtain articles for review.
SourceSearch TermDate Searched
Scopus TITLE-ABS-KEY ((“price spike*” OR “extreme price*” OR “price shock*” OR “price anomal*”) AND (forecast* OR predict*) AND (“deep learning” OR neural OR transformer* OR “long short-term” OR CNN OR LSTM)) 31 January 2026
Web of
Science
TS=(( “price spike*” OR “extreme price*” OR “price shock*” OR “price anomal*”) AND (forecast* OR predict*) AND (“deep learning” OR neural OR transformer* OR “long short-term” OR CNN OR LSTM)) 31 January 2026
IEEE Xplore ((“All Metadata” : “price spike*” OR “All Metadata” : “extreme price*” OR “All Metadata” : “price shock*” OR “All Metadata” : “price anomal*”) AND (“All Metadata” : forecast* OR “All Metadata” : predict*) AND (“All Metadata” : “deep learning” OR “All Metadata” : neural OR “All Metadata” : transformer* OR “All Metadata”: LSTM OR “All Metadata” : CNN OR “All Metadata” : “long short-term”)) 31 January 2026
Table 3. List of papers included in this study.
Table 3. List of papers included in this study.
Market & Ref.Model FamilySpike DefinitionData SourceOutcome
Electricity [3]LSTM HybridThresholdAlberta electricity market (AESO) real-time price data with generation and system variablesImproves multi-day forecast accuracy and enhances spike detection relative to baseline methods.
Electricity [5]Temporal Fusion Transformer (TFT) for quantile forecasting + a separate classification model predicting likelihoodThreshold. “Extreme price” = price > $150ERCOT real-time electricity prices (Houston zone) and NOAA weather dataAchieves lower RMSE (35 vs. 49/42 baselines) and captures extreme prices within the 98th quantile 97% of the time
Electricity [46]DNN spike occurrence classifier. DNN regression + ANN spike calibrationPositive spike: price > μ + 3σPJM electricity market day-ahead price data with operational indicators (e.g., system conditions, demand features)Improves spike detection (F1 ≈ 0.63) and reduces spike-price MAE and sMAPE versus single-stage models without degrading normal-price accuracy.
Electricity [47]Hybrid ML + DL (SVM + LSTM)Statistical threshold (Gaussian 3-σ/2-σ rule)New England electricity market locational marginal price (LMP) data with system variablesImproves forecasting accuracy through explicit spike separation compared to single-model approaches.
Electricity [48]Hybrid: SVM (spike classification) + LSTM (normal prices) + LightGBM (spikes)Statistical threshold: price > μ + 2σ (mean ± 2 standard deviations)Wholesale electricity market LMP data (likely ISO-based, e.g., PJM or similar) with demand, generation, and exogenous variablesSubstantially reduces spike error (MAPE 7.92%, RMSPE 10.84%) with near-perfect spike classification (AUC = 0.9982).
Electricity [6]LSTM (dynamic price prediction) + rule-based threshold detectionThreshold-based: price > predefined electricity or thermal threshold (ρe_th, ρh_th) over next 4 hIndustrial park energy system data with coupled electricity and thermal demand (synthetic or real operational data)Enables rolling 4-h ahead spike warnings based on predicted price thresholds.
Electricity [2]Graph Neural Networks (GNN) for spike detectionFixed thresholds: ≥A$100/MWh (moderate), ≥A$300/MWh (extreme)Australian National Electricity Market (NEM) price data (multivariate time series with system features)Achieves strong spike detection performance (F1 = 0.84) and significantly outperforms LSTM under extreme class imbalance.
Electricity [49]CNN/LSTMFixed extreme price threshold (≈ $150/MWh depending on market)Competitive electricity market price data with engineered features and explainability inputsAchieves very high spike prediction accuracy (~97%) and lower spike-magnitude error compared to baseline models.
Electricity [50]LSTM (DL) + Ensemble Bagging (ML) fusionStatistical threshold (μ + σ; example spike at ≥50 $/MWh) used in post-processingReal-world power system LMP data including load, gas prices, and weather variableIntegrates probabilistic spike estimation into an LMP forecasting pipeline through a dedicated post-processing stage, improving forecasting accuracy under both normal and spike-price conditions (MAE ≈ 0.85 non-spike; ≈ 2.2 spike).
Electricity [51]LLM-based data augmentation + Bayesian NN (MC Dropout)Quantile-based thresholdElectricity market price data (peak price scenarios) with synthetic augmentation via LLMsSignificantly reduces peak-period MAE (e.g., 15.81 vs. 30.32 baseline) and improves arbitrage profitability.
Electricity [52]LLM + CNN-LSTMExplicit extreme threshold + eventsAustralian National Electricity Market (NEM) price data (extreme price forecasting)Substantially reduces extreme price forecasting error, with large MAE improvements during peak events.
Electricity [15]TFTExplicit threshold (≥90th percentile price)Ontario electricity market real-time price, supply, and demand dataReduces MAE and RMSE versus operator benchmarks while achieving higher spike precision (~0.67–0.70) without loss of recall.
Commodity [53]CNN/LSTM/RNN + Attention (neurosymbolic ensemble)Explicit threshold (±2σ rolling)Historical commodity price time series (metals: cobalt, copper, magnesium, nickel)Improves spike detection with +13% F1 and +29% recall over single-model baselines.
Electricity [54]Graph Attention Networks + LSTM encoder–decoder (hybrid GNN–RNN)Adaptive relative thresholds used to label anomaly/spike eventsAustralian National Electricity Market (NEM) multivariate price and system dataOutperforms baseline models in both price accuracy and spike detection across all regions.
Finance [55]Transformer + Wavelet (mWDN-Transformer)price > 2× 28-day moving average = high spike;China stock market limit order book (LOB) high-frequency trading dataImproves precision, recall, and F1 for extreme price movement detection compared to deep learning baselines.
Crypto [4]Feed-forward deep neural networks for quantile/tail-risk predictionQuantile-based tail exceedanceBitcoin (BTC) historical market data including price, technical indicators, and trading featuresShows strong continuous prediction accuracy (R2 ≈ 0.95) and high directional performance (~81%), highlighting a trade-off between accuracy and trading outcomes.
Electricity [56]Multi-output MLP, GRU, CNN-GRU, CNN-LSTMExplicit threshold classification: Class I (<10 AUD), Class II (10–60 AUD), Class III (>60 AUD); approx. 5th & 75th percentileWholesale electricity market price data (ultra-short-term forecasting)Outperforms operator forecasts in ultra-short-term prediction, with substantially higher recall for low-price regimes (~80% vs. ~30%).
Crypto [57]Quantile DNNQuantile-basedCryptocurrency market data (multiple assets, candlestick data + technical indicators)Achieves strong trading performance with high Sharpe ratios (up to 2.73) and improved returns with lower drawdowns.
Electricity [58]Deep Neural Networks (DNN) with Dynamic Sparse TrainingEVT-based extreme price modelling (block maxima) with multiple statistical thresholds; spike defined by magnitude and durationCompetitive electricity market price data (extreme price modelling; likely ISO datasets such as Alberta/NEM)Achieves competitive extreme price forecasting accuracy with strong classification performance across multiple thresholds.
Finance [59]Neural ODEExtreme return/flash crash defined by statistically large price movement and contagion propagation in financial networksFinancial market data (flash crashes and contagion modelling; likely high-frequency asset price time series)Improves detection of extreme financial events by modelling structural dependencies, outperforming baseline methods.
Table 4. Practitioner framework for selecting spike forecasting formulations. The reviewed literature suggests that forecasting objectives, data characteristics, and operational requirements should guide methodological choices before model architecture selection.
Table 4. Practitioner framework for selecting spike forecasting formulations. The reviewed literature suggests that forecasting objectives, data characteristics, and operational requirements should guide methodological choices before model architecture selection.
ConsiderationKey QuestionRecommended Formulation
Forecasting objectiveIs the goal to predict future prices or identify extreme events?Point forecasting for average price prediction; spike forecasting for rare-event detection
Spike labelsAre historical spike events explicitly defined and labelled?Classification or hybrid approaches when labels are available; distributional approaches when labels are unavailable
Uncertainty requirementsIs probability estimation or risk assessment required?Probabilistic classification, quantile forecasting, or distributional forecasting
Spike magnitudeIs the magnitude of extreme events important, or only their occurrence?Hybrid classification–regression or distributional forecasting when magnitude matters
Event rarityHow frequently do spike events occur in the dataset?Imbalance-aware learning, reweighting, oversampling, or specialised loss functions for highly imbalanced problems
Operational applicationHow will forecasts be used in practice?Select evaluation metrics and decision thresholds according to operational objectives and error costs
Data availabilityAre exogenous drivers such as weather, outages, congestion, or market fundamentals available?Incorporate domain-specific features and contextual information where possible
Model selectionWhich model architecture should be used?Select architecture after problem formulation, data characteristics, and evaluation requirements have been established
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sinclair, M.; Shepley, A.J.; Hajati, F. Learning Rare Events: Deep Learning Approaches to Extreme Price Prediction. Forecasting 2026, 8, 52. https://doi.org/10.3390/forecast8030052

AMA Style

Sinclair M, Shepley AJ, Hajati F. Learning Rare Events: Deep Learning Approaches to Extreme Price Prediction. Forecasting. 2026; 8(3):52. https://doi.org/10.3390/forecast8030052

Chicago/Turabian Style

Sinclair, Mark, Andrew J. Shepley, and Farshid Hajati. 2026. "Learning Rare Events: Deep Learning Approaches to Extreme Price Prediction" Forecasting 8, no. 3: 52. https://doi.org/10.3390/forecast8030052

APA Style

Sinclair, M., Shepley, A. J., & Hajati, F. (2026). Learning Rare Events: Deep Learning Approaches to Extreme Price Prediction. Forecasting, 8(3), 52. https://doi.org/10.3390/forecast8030052

Article Metrics

Back to TopTop