1. Introduction
Park-level Integrated Energy Systems (PIESs) enable cascade utilization of electricity, heating, and cooling through the coupling and conversion of multiple energy flows, thereby achieving overall system optimization. Although multi-energy complementarity significantly enhances energy efficiency, it simultaneously introduces a new set of technical challenges. On one hand, the strong coupling among different energy loads makes them difficult to predict independently and accurately; on the other hand, load sequences are influenced by complex interactions among meteorological conditions, equipment operating states, and spatiotemporal dependencies, exhibiting pronounced nonlinear and time-varying characteristics. Consequently, conventional forecasting methods struggle to cope with the combined challenges of high feature dimensionality, intricate coupling mechanisms, and strong nonlinearity. These challenges necessitate the development of new forecasting frameworks capable of deeply integrating temporal features with coupling mechanisms to support efficient operation and refined control of PIESs.
Traditional short-term load forecasting methods primarily include time series analysis and regression techniques. Time series methods utilize intrinsic temporal patterns in historical data for load forecasting. For short-term load forecasting, traditional approaches primarily include time-series analysis and regression-based methods. Time-series models exploit intrinsic temporal patterns within historical data to predict future loads. For example, the Autoregressive Integrated Moving Average (ARIMA) model and its double-seasonal extension (DSARIMA) characterize short-term variations by leveraging the continuity and periodic structure of load sequences [
1,
2]. However, because these models rely on linear assumptions, they often struggle to capture strong nonlinear patterns arising from weather fluctuations, holiday schedules, and user behavior. Their adaptability is further limited when confronted with the intricate coupling mechanisms inherent in multi-energy flow scenarios [
3]. In contrast, machine learning methods—such as Support Vector Machines (SVMs), Artificial Neural Networks (ANNs), and RF—have been widely adopted due to their ability to model nonlinear relationships in short-term load forecasting. For instance, SVMs employ kernel functions to map load data into high-dimensional feature spaces and thereby capture nonlinear dependencies [
4]. Nevertheless, as renewable energy penetration increases, demand-side interactions intensify, and multi-source coupling effects become more pronounced, shallow learning models face growing challenges in capturing long-range temporal dependencies and high-level feature interactions. Consequently, their forecasting accuracy and generalization capability tend to plateau under complex and strongly coupled load conditions [
5,
6]. The rise of deep learning offers new pathways for improving forecasting accuracy. Long Short-Term Memory (LSTM) networks, which alleviate the vanishing gradient problem and effectively capture temporal dependencies, have become a mainstream method. By extracting temporal features from power load sequences, LSTMs have achieved satisfactory results in ultra-short-term load forecasting [
7]. Convolutional Neural Networks (CNNs) also excel in high-dimensional feature extraction; for instance, 1D-CNNs can mine local characteristics in load data and are suitable for load modeling under multiple influencing factors [
8]. Nevertheless, individual deep learning models possess inherent shortcomings: LSTMs are prone to forgetting critical information when input sequences are excessively long [
9]; CNNs are ineffective at modeling long-range dependencies in sequential data [
10]; and model performance depends heavily on hyperparameter tuning and input feature quality, resulting in insufficient robustness [
11]. To overcome the limitations of single models, hybrid models have become a key research focus. One approach uses Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) to decompose load sequences, reconstructs components via sample entropy, and subsequently applies LSTM for forecasting. This method effectively addresses the autocorrelation issue of decomposed components and reduces the RMSE by approximately 30% compared to standard LSTM [
12]. Another hybrid model combining CNNs, LSTM, and an attention mechanism uses CNNs for high-dimensional feature extraction, LSTM to capture temporal patterns, and attention to optimize information weighting. This architecture improved forecasting accuracy by 5.7–7.3% compared to a standalone LSTM in a Combined Heat and Power (CHP) system [
13]. Furthermore, hybrid frameworks integrating stacked ensemble learning and deep reinforcement learning have shown potential for enhancing forecasting stability in distributed IES [
14]. Such hybrid models, by combining the strengths of different algorithms, demonstrate superior performance in modeling nonlinear load characteristics. However, some designs suffer from structural complexity and high computational cost, hindering their application to real-time forecasting in large-scale parks [
15].
While the aforementioned hybrid models focus primarily on improving the forecasting accuracy of individual load sequences, explicitly modeling the coupling relationships among multiple energy flows remains a more fundamental and intrinsic challenge for PIES. Some studies continue to adopt the traditional paradigm of separately modeling cooling, heating, and electricity loads—for example, regional thermal load forecasting based on multiple linear regression [
16] and cooling load prediction using nonlinear autoregressive neural networks [
17]. Although such approaches are methodologically straightforward and easy to implement, they overlook the synergistic interactions and physical constraints inherent among multi-energy flows, often resulting in system-level forecasting deviations. To incorporate energy-flow coupling information, several indirect modeling strategies have been explored. One approach introduces other energy loads as auxiliary inputs through feature engineering—such as incorporating thermal and cooling load data into electricity load forecasting models to implicitly reflect inter-load dependencies [
18]. Another strategy constructs single-task, multi-output models capable of forecasting multiple loads within a unified framework; however, substantial differences in load magnitudes, dynamic behaviors, and variation patterns frequently cause training instability and poor convergence performance [
19]. Multi-Task Learning (MTL) provides a promising alternative for addressing these challenges. By learning shared representations across tasks while preserving task-specific output branches, MTL can explicitly capture coupling mechanisms among multiple loads. For instance, deep multi-task architectures combining Deep Belief Networks (DBNs) with multi-task regression layers achieve an average accuracy improvement of approximately 2.7% over single-task models [
20]. Similarly, integrating MTL with Least-Squares Support Vector Machines (LSSVMs) and optimizing cross-task weights yields an 18.6% enhancement in forecasting accuracy [
21]. In addition, a “shared autoregressive–specific dual-dilated convolutional self-attention” architecture has been proposed, in which a shared AR module extracts linear correlations among turbines while task-specific attention modules capture nonlinear turbine-level features, effectively addressing power forecasting challenges for newly constructed wind farms with limited data [
22]. These studies collectively demonstrate the advantages of MTL in explicitly modeling temporal coupling relationships among multi-energy loads. Nevertheless, existing MTL frameworks often fail to adequately prioritize different load tasks, limiting their ability to satisfy the differentiated forecasting requirements of core loads [
23]. In recent years, more forward-looking deep learning architectures have been widely applied to load forecasting tasks in IES and regional energy networks. Graph Neural Networks (GNNs) and their spatiotemporal extensions—such as ST-GCN and GAT-based models—explicitly leverage the topological structure of park-level energy networks or multi-building clusters, demonstrating strong capability in modeling cross-regional energy flow interactions and consistently outperforming traditional sequence-based models in regional IES load forecasting [
24]. Meanwhile, Transformer-based architectures, benefiting from self-attention mechanisms for long-range temporal dependency modeling, have achieved remarkable progress in tasks such as electricity load forecasting and net-load prediction [
25]. Ensemble frameworks that integrate gradient boosting models with deep neural networks have also shown improved robustness and adaptability in residential demand response and building energy consumption forecasting [
26]. Beyond load forecasting itself, deep learning methods have also been extensively adopted in renewable energy prediction, providing crucial support for developing more sustainable park-level energy systems. For example, hybrid models combining graph convolutional networks and recurrent neural networks have been employed for short-term photovoltaic and wind power forecasting, demonstrating clear advantages in distributed energy scenarios [
27]. Recent studies on hybrid renewable energy systems (HRES) have further integrated LSTM, Bi-LSTM, and MLP models to perform multi-source forecasting of solar, wind, and biomass generation, and have combined such forecasts with load prediction to support energy mix optimization, operational strategy design, and carbon emission assessment [
28]. However, as the variability of renewable energy, the coupling of multi-energy flows, and the complexity of system operation mechanisms continue to increase, these models must handle more diverse and structurally complex input information in practical applications, imposing higher requirements on the forecasting framework.
Whether in advanced hybrid models or MTL frameworks, high-quality feature inputs remain fundamental to achieving reliable predictive performance. As the range of influencing factors continues to expand, load forecasting is increasingly challenged by a rapidly growing feature dimensionality. Meteorological variables, user behavior patterns, equipment operating states, and other load attributes collectively form a high-dimensional input space—particularly in renewable energy scenarios such as photovoltaic (PV) and wind power systems, where irradiance, humidity, wind speed, and cloud cover exhibit strong volatility and complex coupling. Directly feeding such features into forecasting models not only results in substantial computational overhead but may also introduce redundant or noisy variables, ultimately reducing model generalizability [
29]. Traditional feature selection methods exhibit clear limitations. Pearson-correlation-based filtering can only capture linear relationships and is incapable of modeling the more common nonlinear interactions between loads and environmental or behavioral variables [
30]. Experience-driven feature selection is inherently subjective and fails to adapt to evolving system characteristics [
31]. One-time static feature sets also struggle to accommodate the dynamic changes in operational conditions [
32]. These limitations are particularly evident in PV power forecasting; for example, recent findings in solar power plant studies show that different combinations of meteorological inputs can significantly affect forecasting accuracy, and fixed feature sets are often unable to reflect rapidly changing weather conditions [
33]. To address these shortcomings, automated and dynamic feature selection mechanisms have gained increasing attention. The integration of Recursive Feature Elimination (RFE) with deep Bidirectional Long Short-Term Memory (BiLSTM) networks enables continuous updates to feature importance, thereby improving the model’s sensitivity to time-varying factors [
34]. Multi-layer weather classification combined with hierarchical regression has also been developed to reduce redundant features in PV scenarios by structuring inputs according to weather conditions [
35]. Meanwhile, more advanced feature expansion strategies have emerged. Deep spatiotemporal attention models, for instance, explicitly capture temporal and spatial dependencies between loads and meteorological variables, offering greater adaptability to complex multi-energy coupling characteristics than traditional Copula-based approaches [
36]. The fusion of Spatiotemporal Graph Attention Networks (ST-GAT) with multi-stage feature selection further demonstrates strong potential in dynamic feature identification for commercial building cooling load forecasting [
37]. Additionally, Bayesian-optimized LSTM architectures that jointly refine input structures and hyperparameters have shown enhanced robustness and sensitivity to key meteorological features in PV power forecasting [
38]. Despite these advances, achieving an optimal balance between real-time feature selection efficiency and forecasting accuracy in practical PIESs remains an essential challenge for future research.
In summary, while existing research offers diverse methodological approaches for IES load forecasting, several limitations persist for practical park-level applications: hybrid models demonstrate insufficient capability in explicitly addressing multi-energy coupling and dynamic feature selection, while nonlinear components within forecasting residuals remain inadequately compensated. To systematically overcome these challenges, this paper proposes a hybrid forecasting framework that intelligently integrates RFECV, an MTL-LSTM network, and RF. The proposed framework employs a structured three-stage methodology: initially, RFECV performs dynamic feature selection; subsequently, MTL-LSTM simultaneously forecasts multiple load types while capturing their coupling relationships through shared mechanisms; finally, RF is implemented to correct the residuals generated by the primary forecasting model. This modular, phased strategy is designed to significantly enhance both forecasting accuracy and robustness for PIES, while maintaining desirable model interpretability.
4. Results and Discussion
4.1. Hyperparameter Configuration
To thoroughly validate the performance of the proposed hybrid model, this section details the hyperparameter configurations for the developed model and all baseline models. All experiments were conducted under identical hardware and software environments to ensure the comparability of the results.
The MTL-LSTM model was implemented to forecast cooling, heating, and electricity loads using a sliding input window of 24 time steps, enabling the extraction of daily cyclical patterns and short-term temporal dependencies. The shared temporal encoder consisted of two stacked LSTM layers, each with 128 hidden units, followed by a dropout layer with a rate of 0.2 applied before the task-specific output heads. The model was trained using the Adam optimizer with an initial learning rate of 0.001, a batch size of 64, and a maximum of 200 epochs. A dynamic task-weighting mechanism was employed, allowing the relative importance of the three forecasting tasks to be adaptively adjusted throughout the training process.
Training stability and generalization were further enhanced through a collection of regularization and optimization strategies. Gradient clipping was applied to the recurrent layers with a maximum norm of 5 to avoid exploding gradients. An adaptive learning-rate scheduler (ReduceLROnPlateau) was used with a patience of 10 epochs to refine the optimization process when validation loss improvements slowed. Early stopping was employed to terminate training when validation performance failed to improve for 20 consecutive epochs, thereby preventing overfitting. In addition, L2 weight decay with a coefficient of 0.0001 was incorporated into the Adam optimizer to constrain model complexity.
For the second-stage residual correction, an individual RF model was trained for each energy loads. The RF models used the RFECV-selected feature subsets concatenated with the preliminary MTL-LSTM forecasts to learn the residual patterns that the deep network did not fully capture. The primary RF hyperparameters—including the number of trees, maximum depth, and minimum samples per leaf—were determined through grid search combined with a rolling time-series cross-validation procedure. This residual-learning configuration enables the RF to model localized fluctuations and non-stationary behaviors present in the residual sequences.
All experiments were conducted on a personal workstation equipped with a 12th Gen Intel Core i5-12600KF CPU (3.70 GHz), 16 GB of RAM, and an NVIDIA GeForce RTX 4060 Ti GPU with 8 GB of dedicated memory. The implementation was developed in Python 3.9.10, with the deep learning components built using PyTorch 2.0.1 and the RF models implemented via scikit-learn 1.6.1. This hardware–software configuration ensures stable computational performance and reproducibility across all experiments.
4.2. Experimental Results
To comprehensively evaluate the performance of the proposed hybrid RFECV + MTL-LSTM-RF model in park-level energy station load forecasting, a comparative analysis was conducted against four benchmark models: RFECV + MTL-LSTM, Pearson-Correlation-Filtered MTL-LSTM (Pearson + MTL-LSTM), MTL-LSTM, and the conventional LSTM. All models were trained and tested on identical datasets, using the same input features, normalization procedures, and hyperparameter configurations to ensure fair comparability.
Table 2 summarizes the forecasting performance of all models in terms of MAPE, RMSE, and MAE for cooling, heating, and electricity loads. The results demonstrate that the proposed RFECV + MTL-LSTM-RF model consistently outperforms the other approaches across all three load types. Specifically, the multi-task shared-representation structure of the MTL-LSTM model yields markedly better accuracy than the single-task LSTM, confirming that inter-task knowledge transfer effectively enhances forecasting precision. With the incorporation of Pearson-correlation-based feature filtering, the Pearson + MTL-LSTM model achieves further reduction in prediction errors, indicating that removing linearly redundant variables helps lower noise in the input space and improves model expressiveness. When the RFECV feature-selection mechanism is introduced, redundant variables are removed and the input dimensionality is optimized, thereby reducing noise interference caused by irrelevant features. Consequently, the RMSE values for cooling, heating, and electricity loads decrease to 1.428 GJ/h, 2.374 GJ/h, and 1.934 kW, respectively. Finally, by incorporating the RF-based residual-correction module, the forecasting errors are further mitigated, with the RMSE reduced to 1.159 GJ/h, 2.081 GJ/h, and 1.641 kW for the three loads, respectively. These results validate the effectiveness of the model’s modular and staged design.
To further improve the reliability of the evaluation, a nonparametric bootstrap procedure with 1000 resampling iterations was used to estimate the 95% confidence intervals (CIs) of MAPE, RMSE, and MAE for each model. The results show that the RFECV + MTL-LSTM-RF model consistently produces narrower intervals compared with the other models. For example, its cooling-load RMSE CI lies between approximately 1.13 and 1.20 GJ/h, which is clearly tighter than that of the Pearson + MTL-LSTM model (about 1.39–1.48 GJ/h) and much narrower than that of the baseline MTL-LSTM (around 1.57–1.65 GJ/h). Similar patterns of CI contraction are also observed for the heating and electricity loads. These findings indicate that the improvements of the proposed hybrid framework are statistically robust rather than a result of sampling variability.
Figure 6 illustrates the comparison between predicted and actual values for the three energy loads across a representative week. As shown in the figure, the MTL-LSTM model provides a basic level of fit, while the Pearson + MTL-LSTM variant achieves visibly closer alignment with the actual profiles by improving the tracking of peak–valley patterns and local fluctuations. After applying RFECV-based feature refinement, the RFECV + MTL-LSTM model further suppresses noise-induced oscillations and enhances the smoothness and stability of the predicted curves. With the addition of RF-based residual correction, the final RFECV + MTL-LSTM-RF model attains the closest match to the measured loads—particularly during periods of intensive fluctuations and peak demand—where the amplitude of prediction deviation is significantly reduced.
To further confirm that the performance improvement stems from a genuine enhancement in distributional accuracy—rather than spurious fits at isolated points—
Figure 7 presents the daily prediction curves together with the corresponding box-plot distributions of relative errors. In the typical daily load profile, the Pearson + MTL-LSTM model already exhibits visibly better alignment with the actual load trajectory, reducing local biases and capturing peak–valley transitions more accurately. The RFECV + MTL-LSTM model further narrows the deviation band and mitigates short-term overshooting or undershooting, indicating that feature refinement contributes to more stable temporal behavior. Building on these advantages, the RFECV + MTL-LSTM-RF model delivers the closest overall fit to the true load profile throughout the day, with the smallest discrepancies during periods of rapid load variation. The box plots reveal the same progressive improvement. Compared with Pearson + MTL-LSTM, the RFECV + MTL-LSTM model exhibits a noticeably reduced interquartile range (IQR) and shorter whiskers, reflecting contractions in both central dispersion and tail variability. On this basis, the RFECV + MTL-LSTM-RF model achieves the smallest IQR, the shortest whiskers, and the fewest extreme deviations among the three models. Quantitatively, the average relative errors for cooling, heating, and electricity loads decrease from 0.0535, 0.0329, and 0.0433 under Pearson + MTL-LSTM to 0.0503, 0.0299, and 0.0378 under RFECV + MTL-LSTM, and are further reduced to 0.0329, 0.0250, and 0.0279 with RFECV + MTL-LSTM-RF. This contraction in both error magnitude and variance indicates that prediction uncertainty is effectively suppressed, confirming that the model yields more stable and reliable results overall. These findings demonstrate that the residual-correction layer not only mitigates random fluctuations and upper-tail deviations but also achieves a bidirectional optimization of the error distribution, reinforcing the robustness of the final hybrid model.
To further examine the model’s generalization capability and stability, residual scatter analysis was performed for all forecasting tasks, as shown in
Figure 8a–c. In these plots, the horizontal axis represents the predicted values, the vertical axis denotes the residuals, and the standard deviation (σ) is introduced to quantify residual dispersion. Among the baseline models, the Pearson + MTL-LSTM variant already produces slightly more compact and less biased residual clouds than the original MTL-LSTM, reflecting the benefits of removing linearly redundant features. The RFECV + MTL-LSTM model further reduces tail fluctuation and attenuates large-magnitude deviations, suggesting that feature refinement contributes to more stable error behavior. The RFECV + MTL-LSTM-RF model exhibits the most concentrated and symmetric residual distribution around zero, markedly outperforming the baseline models. Its residual points cluster densely near the horizontal axis with visibly fewer extreme departures, indicating consistently lower prediction uncertainty across the entire range of load conditions. Numerically, the residual standard deviations for cooling, heating, and electricity loads are 1 GJ/h, 1.83 GJ/h, and 445.05 kW, respectively—approximately 30% lower on average than those of the Pearson + MTL-LSTM and RFECV + MTL-LSTM models. This hierarchical reduction in σ—from MTL-LSTM to Pearson + MTL-LSTM, then to RFECV + MTL-LSTM, and finally to RFECV + MTL-LSTM-RF—demonstrates the stepwise effectiveness of correlation-based filtering, feature refinement, and residual correction. This confirms that the residual correction module effectively enhances forecasting stability and generalization across varying load magnitudes and operating conditions.
The above results systematically verify the superiority of the proposed framework from the perspectives of forecasting accuracy, residual-distribution characteristics, and generalization performance. Nevertheless, the computational efficiency of the model is equally critical for evaluating its suitability in engineering applications, as practical deployment requires not only accurate predictions but also manageable offline computation and sufficiently fast online forecasting. Assessing these aspects helps to clarify whether the proposed multi-stage architecture can operate effectively within real IES scheduling environments.
Table 3 summarizes the time consumption of each model across the feature-selection, training, residual-correction, and single-step forecasting stages.
The conventional single-task LSTM requires independent training for cooling, heating, and electricity loads, whereas the MTL-LSTM jointly learns all three tasks within a unified temporal representation. This shared-learning mechanism substantially reduces redundant feature extraction and yields a baseline training time of 176.35 s for the multi-task model. Incorporating Pearson-based filtering further decreases the input dimensionality and shortens the training time by approximately 4% relative to the MTL-LSTM baseline, reflecting the benefit of removing linearly redundant predictors and stabilizing the subsequent learning process. RFECV, in turn, provides a more comprehensive feature-refinement mechanism by iteratively eliminating weakly relevant variables under a cross-validation scheme. This procedure reduces the effective feature space before multi-task learning and contributes to improved model conditioning. Although RFECV requires 267.82 s for feature selection and increases the total offline cost relative to the MTL baseline, the RFECV-enhanced hybrid models deliver notable gains in forecasting accuracy and exhibit consistently improved residual stability, as evidenced in
Table 2. These performance improvements provide a clear justification for the additional offline computation. Moreover, the RFECV + MTL-LSTM-RF model maintains a forecasting latency that is slightly faster than that of the MTL-LSTM baseline, indicating that the enhanced model retains efficient inference capability despite its more elaborate training pipeline.
Taken as a whole, the analyses of forecasting accuracy, residual-distribution characteristics, and computational efficiency presented in this section consistently demonstrate that the proposed hybrid framework delivers reliable predictive performance across all load types while maintaining stable error behavior under both typical and volatile conditions. These results underscore the universality and robustness of the multi-layer integrated architecture, indicating that the gains in accuracy and stability are achieved with a computational cost that remains acceptable for practical IES applications. Building on these empirical findings, the following discussion further examines the underlying mechanisms that contribute to the effectiveness of the proposed approach.
4.3. Discussion
The experimental results provide several insights into the mechanisms underlying the performance of the proposed hybrid RFECV + MTL-LSTM-RF framework. The multi-task architecture enables the model to leverage shared temporal patterns among cooling, heating, and electricity loads, which often respond jointly to weather and operational conditions. This shared representation helps the model capture synchronized fluctuations that single-task approaches tend to miss.
The RFECV-based feature refinement further stabilizes the forecasting process by removing redundant or weakly relevant features, thereby reducing noise transmission into the learning pipeline. This effect is particularly evident when comparing Pearson + MTL-LSTM with RFECV-enhanced models, suggesting that systematic data-driven feature screening provides more consistent benefits than simple correlation-based filtering alone. Additionally, the RF residual-correction layer effectively addresses nonlinear and heteroscedastic components of forecasting errors, yielding more compact residual distributions and improving robustness during periods of rapid load variation.
These observations also have implications for sustainable IES. As the proportion of renewable resources such as photovoltaics and wind power continues to grow, reliable short-term load forecasting becomes increasingly important for maintaining operational stability and supporting flexible scheduling. The modular structure of the proposed framework makes it well suited for future extension toward multi-energy prediction and renewable-integration scenarios.
While the proposed framework demonstrates strong predictive capability, it still operates as an open-loop forecasting model under assumed system conditions and does not explicitly account for the physical and economic constraints inherent to real IES operation. Physically, equipment capacity limits, ramping capabilities, and dynamic response characteristics define the feasible region of short-term load variations, meaning that actual loads cannot freely follow weather-driven signals or historical inertia. Economically, time-of-use tariffs, demand-response incentives, and renewable-driven dispatch adjustments reshape load profiles through price- and cost-driven behavioral changes. Incorporating such physical and economic constraints into future forecasting frameworks would shift the task from predicting the most likely load trajectory to generating feasible trajectories that reflect both operational limitations and market-responsive behavior.
Nonetheless, the discussion should acknowledge that the model has not yet incorporated renewable generation dynamics, energy storage behavior, or cross-site validation. These aspects are essential for fully supporting next-generation sustainable energy station operation and represent important directions for future work.
Looking forward, future integration of renewable energy forecasting will enable the proposed framework to operate in a more unified multi-energy setting. In such scenarios, RFECV can be extended to simultaneously select informative predictors for both load-side and generation-side variables—for example, irradiance-, cloud-cover-, and temperature-related features for PV generation, and wind-speed-, turbulence-, and direction-related features for wind power. This unified feature-selection process would allow the model to identify cross-dependencies not only among cooling, heating, and electricity loads, but also between loads and renewable outputs. Likewise, embedding renewable generation forecasting tasks into the MTL-LSTM architecture would allow the shared temporal encoder to learn broader multi-energy interactions, capturing correlations such as load–PV coupling under high-irradiance conditions or wind–electricity complementarity during rapid weather changes. Such an extension would further enhance the framework’s applicability to renewable-rich IES and support integrated forecasting and scheduling needs in future sustainable energy stations.
5. Conclusions
This study presents a hybrid RFECV + MTL-LSTM-RF forecasting framework for cooling, heating, and electricity loads in park-level energy stations. By integrating multi-task temporal representation learning, data-driven feature refinement, and residual-based error correction, the proposed framework achieves significant and consistent accuracy improvements over benchmark models. The hybrid approach benefits from the complementary strengths of its components: shared representations across correlated energy loads, optimized and noise-reduced feature inputs, and an additional correction layer capable of capturing complex residual structures, while the overall computational cost of the hybrid architecture remains acceptable and its forecasting latency stays well within the sub-second range, supporting real-time applicability in practical IES operations.
Methodologically, the work contributes a structured and interpretable modeling pipeline, in which feature selection, sequence prediction, and residual learning are tightly coupled. The RFECV procedure effectively removes redundant meteorological and operational variables, thereby enhancing learning stability. The MTL-LSTM architecture exploits cross-load dependencies to better capture joint fluctuations and common temporal patterns. Finally, the RF-based residual module addresses nonlinear and heteroscedastic error components that conventional neural networks may overlook, leading to improved distributional fidelity and reduced extreme deviations.
Beyond its methodological contributions, the proposed framework holds practical value for sustainable IES. Accurate short-term load prediction supports refined scheduling, reduced auxiliary energy consumption, and improved utilization of distributed resources. As renewable energy penetration continues to increase—particularly photovoltaics and wind power—the hybrid architecture offers a scalable basis for integrating renewable generation forecasting and multi-energy coordination. The enhanced robustness and reduced uncertainty achieved by the framework contribute directly to the efficient and low-carbon operation of modern park-level energy stations.
Several limitations should be acknowledged. The analysis is based on data from a single energy station, and broader validation across multiple sites with diverse climatic and operational conditions is needed. The current framework focuses solely on load-side forecasting and does not incorporate renewable outputs, energy storage behavior, or cross-site generalization. Moreover, physical constraints—such as equipment capacity limits, ramping capabilities, and dynamic response characteristics—and economic dispatch factors, including time-of-use pricing and demand-response incentives, were not explicitly modeled. Future work will extend the framework toward holistic multi-energy forecasting by integrating renewable generation prediction and storage dynamics. Such an extension would enable RFECV to conduct unified feature selection over both load-side and generation-side variables, while allowing the MTL-LSTM architecture to learn broader cross-energy dependencies. Together, these developments will further enhance the practicality and scalability of the proposed method for real-world sustainable and resilient energy systems.