EEMD-TiDE-Based Passenger Flow Prediction for Urban Rail Transit
Abstract
1. Introduction
- (1)
- Architectural synergy: TiDE’s channel-independent assumption (transforming multivariate forecasting into parallel univariate tasks) perfectly aligns with EEMD’s IMF components (each a stationary time series). Its residual block (with dropout, ReLU, residual connections, and LayerNorm) further optimizes feature extraction from decomposed signals.
- (2)
- Computational efficiency: EEMD’s O(N2) complexity (N = sequence length) and TiDE’s O(L·H) forecasting complexity (L = lookback window, H = horizon) enable high-accuracy predictions without processing raw non-stationary data, preserving efficiency while enhancing performance.
- (3)
- Hierarchical covariate utilization: TiDE’s unified treatment of static/dynamic covariates is synergized with EEMD’s multi-resolution decomposition, enabling a differential covariate–IMF association analysis for refined feature learning.
2. Methodology
2.1. Decomposition: EEMD
2.2. TiDE
2.3. EEMD-TiDE
3. Experiments
3.1. Dataset
3.2. Model Configurations and Evaluation Metrics
3.3. Baseline Models
- ARIMA: The Autoregressive Integrated Moving Average (ARIMA) model is a classical linear method used for time series forecasting. It is defined by three key parameters: the autoregressive order (pp), the differencing order (dd), and the moving average order (qq). Optimal parameter selection is achieved using the Akaike Information Criterion (AIC), balancing predictive accuracy and model complexity. Its parameters are configured as follows: the lag order is set to 2, the degree of difference is 1, and the order of moving averages is also 1.
- LSTM: Long Short-Term Memory, a variant of recurrent neural networks (RNNs), employs collaborative mechanisms of input, forget, and output gates to effectively capture long-term dependencies in sequential data. This architecture mitigates gradient vanishing issues, enabling the stable processing of temporal sequences and achieving high precision in complex forecasting tasks. In this model, there is one fully connected layer accompanied by two hidden layers. The hidden states are configured with a dimension of 128, while the fully connected layer comprises 64 neurons.
- GRU: The Gated Recurrent Unit simplifies the LSTM architecture by merging the input and forget gates into a single update gate and eliminating the separate cell state. This structural optimization reduces both the parameter count and computational complexity while preserving the ability to model long-term dependencies, thereby enhancing training efficiency and inference speed. In this model, there is one fully connected layer accompanied by two hidden layers. The hidden states are configured with a dimension of 128, while the fully connected layer comprises 64 neurons.
- TCN: The Temporal Convolutional Network utilizes dilated convolutional layers to parallelize temporal data processing, enabling the efficient capture of long-range dependencies through stacked convolutional structures. Unlike traditional RNNs, the TCN avoids sequential computation bottlenecks, offering superior computational efficiency for large-scale time series forecasting. In this model, there is one fully connected layer accompanied by two hidden layers. The hidden states are configured with a dimension of 128, while the fully connected layer comprises 64 neurons. Meanwhile, the number of channels of the TCN is [64, 128, 256], and the convolutional kernel size is 3.
- Transformer: The Transformer architecture leverages an encoder–decoder framework combined with self-attention mechanisms to dynamically weight the importance of different temporal positions in input sequences. This adaptive weighting mechanism enables the robust modeling of long-term dependencies and exceptional contextual understandings, making it highly effective for the high-precision forecasting of complex temporal patterns. In this model, the conventional Transformer is composed of three encoder layers and three decoder layers, each equipped with eight multi-head attention mechanisms. The dimensions for dk and dv are both configured to be 64. Following the Transformer’s output, it is fed into two fully connected layers, the first with 128 neurons and the second with 64 neurons.
- Informer: The Informer is an innovative variant of the Transformer model, specifically designed for long-sequence time series prediction tasks. It introduces a sparse probabilistic self-attention mechanism, which significantly reduces the computational complexity from O(L2) to O(L log L), effectively minimizing both the computation and memory usage during the processing of extended sequences. Furthermore, the Informer employs a self-attention distillation technique that optimizes the allocation of computational resources and memory consumption for long sequences by emphasizing the most salient attentional patterns. This approach contrasts with traditional step-by-step decoding methods, as the Informer’s generative decoder can predict entire sequences in a single pass, thereby markedly enhancing the inference speed. These enhancements enable the Informer to perform better than existing methods across multiple large-scale datasets. The parameters of the Informer are the same as the Transformer.
3.4. Prediction Performance
4. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Li, M.; Zhang, C. An Urban Metro Section Flow Forecasting Method Combining Time Series Decomposition and a Generative Adversarial Network. Sustainability 2024, 16, 607. [Google Scholar] [CrossRef] [Scilit]
- Owais, M.; Allam, A.A. Adaptive Optimization of Traffic Sensor Locations Under Uncertainty Using Flow-Constrained Inference. Appl. Sci. 2025, 15, 10257. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Kang, Y.; Hyndman, R.J.; Li, F. Distributed ARIMA models for ultra-long time series. Int. J. Forecast. 2023, 39, 1163–1184. [Google Scholar]
- Zhao, J.; Yu, Z.; Yang, X.; Gao, Z.; Liu, W. Short term traffic flow prediction of expressway service area based on STL-OMS. Phys. A Stat. Mech. Its Appl. 2022, 595, 126937. [Google Scholar] [CrossRef] [Scilit]
- Lin, G.; Lin, A.; Gu, D. Using support vector regression and K-nearest neighbors for short-term traffic flow prediction based on maximal information coefficient. Inf. Sci. 2022, 608, 517–531. [Google Scholar] [CrossRef] [Scilit]
- Yu, B.; Lee, Y.; Sohn, K. Forecasting road traffic speeds by considering area-wide spatio-temporal dependencies based on a graph convolutional neural network (GCN). Transp. Res. Part C Emerg. Technol. 2020, 114, 189–204. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.-F.; Chen, Y.-P.P. Hybrid deep learning and empirical mode decomposition model for time series applications. Expert Syst. Appl. 2019, 120, 128–138. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Wu, F.; Liu, Z.; Wang, K.; Wang, F.; Qu, X. Can language models be used for real-world urban-delivery route optimization? Innovation 2023, 4, 100520. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tselentis, D.I.; Vlahogianni, E.I.; Karlaftis, M.G. Improving short-term traffic forecasts: To combine models or not to combine? IET Intell. Transp. Syst. 2015, 9, 193–201. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Zhang, Y.; Haghani, A. A hybrid short-term traffic flow forecasting method based on spectral analysis and statistical volatility model. Transp. Res. Part C Emerg. Technol. 2014, 43, 65–78. [Google Scholar] [CrossRef] [Scilit]
- Pan, Y.A.; Guo, J.; Chen, Y.; Cheng, Q.; Li, W.; Liu, Y. A fundamental diagram-based hybrid framework for traffic flow estimation and prediction by combining a Markovian model with deep learning. Expert Syst. Appl. 2024, 238, 122219. [Google Scholar] [CrossRef] [Scilit]
- Deng, S.; Du, J.; Zhang, J.; Wang, X. Deep learning approach for short-term entry passenger flow forecasting in urban rail transit stations. Eng. Appl. Artif. Intell. 2026, 163, 112989. [Google Scholar]
- Mao, Y.; Yu, X. A hybrid forecasting approach for China’s national carbon emission allowance prices with balanced accuracy and interpretability. J. Environ. Manag. 2023, 351, 119873. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, H. Introduction to the Hilbert–Huang Transform And Its Related Mathmatical L Problems. In Hilbert-Huang Transform and Its Applications; World Scientific Publishing: Singapore, 2003; pp. 1–26. [Google Scholar]
- Wei, Y.; Chen, M.-C. Forecasting the short-term metro passenger flow with empirical mode decomposition and neural networks. Transp. Res. Part C Emerg. Technol. 2012, 21, 148–162. [Google Scholar] [CrossRef] [Scilit]
- Yu, X.; Shi, S.; Xu, L. A spatial–temporal graph attention network approach for air temperature forecasting. Appl. Soft Comput. 2021, 113, 107888. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Chen, Y.; Tso, G.K.F. Price forecasting in the precious metal market: A multivariate EMD denoising approach. Resour. Policy 2017, 54, 9–24. [Google Scholar] [CrossRef] [Scilit]
- Cao, J.; Li, Z.; Li, J. Financial time series forecasting model based on CEEMDAN and LSTM. Phys. A Stat. Mech. Its Appl. 2019, 519, 127–139. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.; Hasan, N.; Deng, C.; Bao, Y. Multivariate empirical mode decomposition based hybrid model for day-ahead peak load forecasting. Energy 2022, 239, 122245. [Google Scholar] [CrossRef] [Scilit]
- Hao, W.; Sun, X.; Wang, C.; Chen, H.; Huang, L. A hybrid EMD-LSTM model for non-stationary wave prediction in offshore China. Ocean Eng. 2022, 246, 110566. [Google Scholar] [CrossRef] [Scilit]






| Categories | Version |
|---|---|
| Central Processing Unit (CPU) | 13th Gen Intel(R) Core(TM) i7-13620H 2.40 GHz |
| Operating System (OS) | Windows 11 Professional x64 |
| Random Access Memory (RAM) | 16.0 GB |
| Graphics Processing Unit (GPU) | NVIDIA GeForce GTX 4060 8 GB |
| Programming Language | Python 3.9.18 |
| Development Environment | VSCode 1.84.1 |
| CUDA 12.1 | |
| Main Python Modules | PyTorch 2.3.0 |
| EMD-signal 1.4.0 | |
| NumPy 1.26.4 | |
| Pandas 1.5.3 | |
| Matplotlib 3.8.3 |
| Model | MAE | RMSE | MAPE |
|---|---|---|---|
| ARIMA | 8.549 | 15.783 | 7.447 |
| LSTM | 5.364 | 9.917 | 4.672 |
| GRU | 5.237 | 9.328 | 4.391 |
| TCN | 4.948 | 8.972 | 4.313 |
| Transformer | 4.895 | 8.882 | 4.294 |
| EEMD–Transformer | 4.801 | 8.733 | 4.203 |
| Informer | 5.085 | 9.693 | 4.413 |
| TiDE | 4.732 | 8.612 | 4.118 |
| EEMD-TiDE | 4.513 | 7.951 | 4.003 |
| Model | MAE | RMSE | MAPE |
|---|---|---|---|
| ARIMA and LSTM | 0.0003 | 0.0472 | 0.0124 |
| ARIMA and GRU | 0.0283 | 0.0013 | 0.0101 |
| ARIMA and TCN | 0.0482 | 0.0456 | 0.0283 |
| ARIMA and Transformer | 0.0143 | 0.0421 | 0.0184 |
| ARIMA and EEMD–Transformer | 0.0143 | 0.0335 | 0.0471 |
| ARIMA and Informer | 0.0352 | 0.0105 | 0.0384 |
| ARIMA and TiDE | 0.0402 | 0.0173 | 0.0205 |
| ARIMA and EEMD-TiDE | 0.0092 | 0.0076 | 0.0117 |
| LSTM and GRU | 0.0491 | 0.0702 | 0.0165 |
| LSTM and TCN | 0.0463 | 0.0278 | 0.0089 |
| LSTM and Transformer | 0.0415 | 0.0194 | 0.0247 |
| LSTM and EEMD–Transformer | 0.0389 | 0.0065 | 0.0423 |
| LSTM and Informer | 0.0152 | 0.0342 | 0.0480 |
| LSTM and TiDE | 0.0686 | 0.0113 | 0.0364 |
| LSTM and EEMD-TiDE | 0.0256 | 0.0172 | 0.0375 |
| GRU and TCN | 0.0453 | 0.0231 | 0.0124 |
| GRU and Transformer | 0.0397 | 0.0264 | 0.0183 |
| GRU and EEMD–Transformer | 0.0352 | 0.0430 | 0.0216 |
| GRU and Informer | 0.0473 | 0.0331 | 0.0147 |
| GRU and TiDE | 0.0442 | 0.0263 | 0.0078 |
| GRU and EEMD-TiDE | 0.0367 | 0.0403 | 0.0176 |
| TCN and Transformer | 0.0235 | 0.0372 | 0.0059 |
| TCN and EEMD–Transformer | 0.0414 | 0.0245 | 0.0138 |
| TCN and Informer | 0.0146 | 0.0335 | 0.0475 |
| TCN and TiDE | 0.0276 | 0.0108 | 0.0358 |
| TCN and EEMD-TiDE | 0.0472 | 0.0308 | 0.0124 |
| Transformer and EEMD–Transformer | 0.0436 | 0.0270 | 0.0080 |
| Transformer and Informer | 0.0363 | 0.0419 | 0.0485 |
| Transformer and TiDE | 0.0297 | 0.0153 | 0.0451 |
| Transformer and EEMD-TiDE | 0.0264 | 0.0079 | 0.0374 |
| EEMD–Transformer and Informer | 0.0409 | 0.0469 | 0.0323 |
| EEMD–Transformer and TiDE | 0.0142 | 0.0428 | 0.0255 |
| EEMD–Transformer and EEMD-TiDE | 0.0071 | 0.0357 | 0.0396 |
| Informer and TiDE | 0.0242 | 0.0162 | 0.0336 |
| Informer and EEMD-TiDE | 0.0410 | 0.0195 | 0.0218 |
| TiDE and EEMD-TiDE | 0.0105 | 0.0332 | 0.0268 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cheng, D.; Zhang, Y.; Li, H. EEMD-TiDE-Based Passenger Flow Prediction for Urban Rail Transit. Electronics 2026, 15, 529. https://doi.org/10.3390/electronics15030529
Cheng D, Zhang Y, Li H. EEMD-TiDE-Based Passenger Flow Prediction for Urban Rail Transit. Electronics. 2026; 15(3):529. https://doi.org/10.3390/electronics15030529
Chicago/Turabian StyleCheng, Dongcai, Yuheng Zhang, and Haijun Li. 2026. "EEMD-TiDE-Based Passenger Flow Prediction for Urban Rail Transit" Electronics 15, no. 3: 529. https://doi.org/10.3390/electronics15030529
APA StyleCheng, D., Zhang, Y., & Li, H. (2026). EEMD-TiDE-Based Passenger Flow Prediction for Urban Rail Transit. Electronics, 15(3), 529. https://doi.org/10.3390/electronics15030529

