1. Introduction
In recent years, rapid urban expansion and increasing travel demands have resulted in severe traffic congestion in cities [
1]. High-capacity and high-speed urban rail transit can provide comfortable and punctual service and has become an efficient solution to reduce urban traffic pressure [
2]. In mega-metro systems such as Seoul’s subway, peak-period passenger flows often lead to severe overcrowding in transfer stations [
3]. By the end of 2025, China’s urban rail transit systems consisted of 343 lines with a total length of 11,710.3 km, accommodating 33.24 billion passenger trips annually—a 3.1% year-on-year increase. Empirical studies also show that in rapidly urbanizing regions, average daily ridership intensity can exceed 3.6 million passengers per day on major lines, reflecting both network dependence and growth pressure on the metro infrastructure [
4]. However, with the increases in passenger flows, urban rail systems are encountering several operational issues, including overcrowding during peak hours, slow passenger movement on the platforms and increasing safety hazards. All of these challenges highlight the demand for accurate forecasts of passenger flow in urban rail systems.
Passenger flow indicates the change in passenger numbers in urban rail transit stations and lines over time. An accurate forecast of passenger flow enables informed decision-making in the operation, management, and planning of an urban rail station, as well as the entire line (e.g., efficient passenger flow management of stations and train scheduling on the line). Therefore, forecasting plays an important role in improving the efficiency and reliability of urban transportation systems [
5], benefitting people’s lives and sustainable city development.
The accuracy of forecasting passenger flows in urban rail transit at both the station and line level depends on the development of an effective model. However, passenger flow exhibits complex characteristics, such as nonlinearity and uncertainty, since it is influenced by multiple sources of information (e.g., daily commute patterns and weather conditions). Such characteristics make modeling the passenger flow in urban rail transit difficult. In this case, there is a great need for a forecasting model that can effectively capture the complexity of passenger flow [
6,
7,
8]. Moreover, recent research has found that passenger flow forecasting remains challenging due to the highly nonlinear and stochastic nature of ridership dynamics, which are significantly influenced by temporal and spatial factors across urban rail networks [
9].
2. Literature Review
2.1. Urban Expansion and Traffic Congestion
Rapid urbanization is reshaping travel demand and is a key cause of chronic congestion in metropolitan regions [
10]. Liu et al. [
11] argued that relieving this pressure requires integrated strategies—transit-oriented development, multimodal integration, and prioritizing public transit—to shift trips away from private cars. Zhou et al. [
12] used a bi-level programming model to show that the location of transit-oriented development stations strongly affects travel behavior, multimodal network equilibrium, and thus congestion. Serdar Dindar [
13] noted that passenger flows are highly complex, reflecting urban activities and spatio-temporal heterogeneity rather than just commute patterns or weather, while Wang et al. [
14] found that after-work activities—residential, leisure, and secondary destinations—strongly shape metro arrival times, usage patterns, and peak loads. Gao et al. [
15] showed that spatio-temporal variability and differences across population groups are crucial for assessing how land use and transport affect mobility. The rise of shared micromobility, when integrated with transit, further complicates passenger behavior and undermines traditional planning assumptions [
16]. Panel data also indicate that as urban rail networks grow more structurally complex, the coordination between population distribution and line layout has a major impact on passenger flow dynamics [
17].
Collectively, these studies show that urban traffic congestion is not just a consequence of a supply–demand imbalance but is a multifaceted phenomenon shaped by urban form, land use policies, multimodal integration, and changing travel behaviors. The success of mitigation strategies—such as transit-oriented development and shared mobility—depends on understanding these dynamics. This insight motivates an examination of how passenger flow characteristics reflect these complex urban interactions. In the following subsection, we review the literature on passenger flow characteristics and their determinants, providing the analytical basis for robust forecasting frameworks.
2.2. Passenger Flow Characteristics and Influencing Factors
Passenger flow characteristics involve complex interactions among passenger behavior, network structure, and the built environment. Yuan et al. [
18] reviewed urban rail passenger flow assignment methods, identifying an evolution from graph-theoretic path assignment to advanced frameworks that account for travel heterogeneity, dynamic passenger–train matching, and AI-based approaches, thereby improving our understanding of flow distribution. Zeng et al. [
19] used random forest models to study metro–bus transfers in Shanghai and found that the transfer time and bus routes were dominant influencers of peak-hour transfer behavior, while land-use mix was more influential off-peak, underscoring temporal differences in intermodal integration. Chen et al. [
20] applied a multi-scale spatio-temporal geographically weighted regression model with 13 POI types using demographic data for Hangzhou, showing that business and residential variables strongly increased morning peak inbound flows, while scenic spots suppressed them and corporate variables had non-stationary effects—boosting weekday but reducing weekend passenger flow.
Recent empirical studies have further clarified how multiple factors shape passenger demand across space and time. Zhao et al. [
21] used Lasso-multiscale geographically weighted regression on Hangzhou AFC data and found that the ultra-peak boarding time was mainly influenced by transportation, residential, and working land use and closeness centrality, whereas the ultra-peak boarding volume was shaped by residential, working, and healthcare land use, bus stop density, and transfer station status. All effects showed strong spatial heterogeneity between urban cores and suburbs. Liu et al. [
22], using the GeoDetector model for Shanghai, showed that interactions between built environment variables and land-use diversity significantly boosted metro ridership, with food and beverage services most influential on weekdays and accommodation services most influential on weekends. Liu et al. [
23] applied panel data methods in Beijing and confirmed the temporal heterogeneity of land-use effects: land-use density increases the number of commuters in the morning peak and generates commuter trips in the evening peak, while land-use diversity promotes trips over longer periods on non-weekdays. Zhang et al. [
24] used an XGBoost-SHAP model with Tianjin metro data and found that diverse travel-related factors accounted for 59.8% and 61.3% of importance on weekends and holidays, respectively, and that residential land strongly interacted with office, shopping, leisure, and tourist attractions across different periods.
Collectively, these studies show that passenger flow is affected by an interplay of temporal patterns, spatial dependencies, land use, transfer behavior, and evolving passenger activities. The documented spatio-temporal heterogeneity—across time periods, station types, and network contexts, including varying built-environment effects, nonlinear thresholds, and strong interaction effects—highlights the need for forecasting models that capture both local variation and system-wide dependencies.
2.3. Forecasting Models
In response to the growing need for accurate passenger flow predictions, there is a substantial body of research devoted to developing forecasting models for urban rail transit. The existing approaches can be broadly categorized into two main classes—statistical models and machine learning methods—each with distinct strengths and limitations in capturing the underlying dynamics of passenger demand.
Common statistical forecasting models are autoregressive integrated moving average (ARIMA) models, seasonal ARIMA (SARIMA) models, state-space models, and gray models. Zhang et al. [
25] employed an enhanced ARIMA model to forecast the passenger flow at Tianfu Square Station on the Chengdu urban rail transit network. Li et al. [
26] extended the ARIMA—into SARIMA—by incorporating seasonal components to achieve a more effective model of periodically varying time series, and used the proposed model to forecast passenger flow on the Guangzhou-to-Zhuhai intercity railway. Liang et al. [
27] integrated state-space models with the k-nearest neighbors algorithm to improve the accuracy of passenger flow forecasts for Changchun rail stations. Song et al. [
28] also successfully employed a fractional-order gray model to forecast short-term traffic flows in the UK.
Although commonly used to forecast passenger flows in urban rail transit, the above statistical models are generally incapable of characterizing the nonlinearity and dynamics in passenger flow data, which reduces their accuracy and adaptability [
29]. As an alternative, machine learning techniques have been extensively adopted to forecast passenger flows in urban rail transit. Due to their strong nonlinear approximation capability, these methods can be used to automatically extract the complex latent patterns and dynamic dependencies in passenger flow data. Therefore, compared to statistical models, they can better address the obvious uncertainty and nonlinearity of real-world passenger flow dynamics. Sun et al. [
30] utilized a support vector machine (SVM) model to forecast passenger flows on the Beijing metro, developing an improved Categorical Boosting (CatBoost) model to forecast the monthly ridership [
31]. Chuwang et al. [
5] used several machine learning models, including XGBoost and AdaBoost, to estimate historical passenger flows in the Changsha urban rail transit network. Wang et al. [
32] proposed an ensemble forecasting framework that incorporates a random forest model to better characterize the heterogeneity of urban rail transit stations in order to estimate the passenger flows at metro stations in Hangzhou.
As a core subdomain of machine learning, deep learning has also been widely employed in urban rail transit passenger flow forecasting considering its strong capability in automated feature extraction and sequence modeling. Among the deep learning methods, the long short-term memory (LSTM) network, an extension of recurrent neural networks (RNNs), is designed to capture the long-range temporal dependencies of time series and has been studied by many researchers. Lu et al. [
33] developed an enhanced LSTM-based model to forecast passenger flows in the Shanghai rail transit network and demonstrated its performance by comparing it with the single LSTM model. Wang et al. [
34] proposed a hybrid forecasting approach that combines LSTM and a bidirectional gated recurrent unit (BiGRU), verifying that the proposed hybrid approach can achieve a higher forecasting accuracy than that achieved using only LSTM or BiGRU. Furthermore, Liu et al. [
7] used the bidirectional LSTM (BiLSTM) network—which can exploit the comprehensive bidirectional contextual information in time-series data—to forecast the passenger flows in the Hangzhou rail transit system. Deng et al. [
35] employed a temporal convolutional network (TCN) to forecast passenger flows in the Suzhou metro. However, these approaches rely heavily on the temporal features of urban rail transit passenger flows to carry out forecasting while neglecting the mutual interdependency and spatial correlations existing among the stations in real-world urban rail transit systems [
36], worsening their performance in passenger flow forecasting.
To address this weakness, recent studies have incorporated spatial information into forecasting models. Graph convolutional networks (GCNs) designed to address graph-structured data can effectively capture the spatial dependencies among the nodes in a graph and have thus been widely used in passenger flow forecasting for urban rail transit. Zhang et al. [
37] developed a multi-GCN with gated recurrent units to model both the spatial and temporal correlations of metro passenger flows in Hangzhou and Shanghai. Zeng et al. [
38] combined GCNs, attention mechanisms, and LSTM networks to forecast passenger flows in the Shenzhen and Hangzhou rail transit systems; in these models, the travel behavior and comprehensive spatio-temporal dependencies are more tightly coupled with the forecasting. Liu et al. [
39] proposed a Dynamic Spatio-Temporal Graph Fusion Network (DSTGFN) to deeply extract and exploit the spatio-temporal patterns of the passenger flows in urban rail transit and applied it to the networks in Hangzhou and Shanghai. Zeng et al. [
40] introduced an Adaptive Spatio-Temporal Hierarchical Network (ASTHN) for passenger flow forecasting, in which a spatio-temporal gated convolutional layer (ST-Gconv) that integrates dilated convolutions, gating mechanisms and adaptive graph convolutions was introduced to enhance the ability to capture the complex spatio-temporal dependencies in passenger flow data. Yin et al. [
41] applied fast Fourier transform to identify latent periodicities in passenger flow time-series and incorporated attention mechanisms with adaptive graph convolutional networks to learn temporal dependencies and spatial patterns separately under distinct periodic regimes. Xie et al. [
42] developed a DSTGFN with alternately stacked multi-time attention modules and verified its improved accuracy in forecasting the passenger flows of urban rail systems in Beijing, Shanghai, and Hangzhou.
2.4. Research Gaps and Contributions
Existing studies have made significant contributions to passenger flow forecasting in urban rail transit stations. However, there are still weaknesses, as shown below.
(1) Although deep learning methods (especially GCNs) have been utilized in existing studies, these studies have not fully explored the scope and granularity of spatio-temporal feature fusion. Moreover, the majority of these works rely on graph neural networks to capture static or dynamic spatial dependencies among neighboring stations; however, they have not used multi-level aggregated information or multi-scale periodic patterns in forecasting. This incomplete feature engineering results in a model that lacks the ability to learn complex dynamics, which restricts forecasting accuracy.
(2) Deep neural networks and their variants require large training datasets and precise hyperparameter tuning to achieve acceptable performance, and their generalization is reduced when the data distribution changes. On the contrary, decision tree-based ensemble models (e.g., random forests and gradient boosting) are typically more robust to noisy features and are more interpretable. Although some studies have combined heterogeneous models to enhance forecasting accuracy, most fusion strategies still rely on simple stacking or a combination based on fixed weights. Such strategies cannot reflect the evolving passenger flow patterns or the spatial structure of the urban rail transit network. As a result, these static fusion strategies cannot adaptively integrate the models, which reduces the forecasting accuracy.
(3) Existing studies did not simultaneously forecast station- and line-level passenger flows. Aggregating station-level forecasts to the line level allows for station-specific errors to transfer and accumulate and further leads to deviations in line-level estimations, leading to large discrepancies between the estimated and actual line-level passenger flows. At the same time, models that only forecast line-level passenger flows cannot provide station-level results. Joint forecasting can provide accurate passenger flows in both urban rail transit stations and lines, which enables better operational, management and planning decisions. However, this has not been highlighted in the literature.
To address these issues, in this study, a passenger flow forecasting model for urban rail transit is proposed that combines multi-source spatio-temporal features and optimized ensemble learning. This study address deficiencies in the current research; its main contributions are as follows.
(1) To address inadequate feature fusion, we propose a multi-source spatio-temporal feature engineering framework to mine the multi-level heterogeneous space–time features of passenger flows in urban rail transit. This framework uses graph-based station adjacency and attention mechanisms to determine inter-station dependencies, clustering stations based on passenger flows. Furthermore, it incorporates diverse time variables (e.g., daily and weekly cycles) and lagged terms to capture the autocorrelation in passenger flow time-series.
(2) A dynamically weighted ensemble forecasting framework with high computational efficiency, a low training cost and improved numerical stability is proposed, in which ExtraTrees and LightGBM—models suitable for structured data—are used as base learners to calculate preliminary forecasts. A particle swarm optimization (PSO) algorithm is also adopted to determine their optimal aggregation weights to obtain the station-level forecasting results. This framework preserves the stability and interpretability of ensemble learning, and the contributions of the base learners can be adaptively adjusted to address the varying data distributions. Therefore, the overall accuracy and robustness of forecasting can be improved.
(3) To address the inconsistencies between station- and line-level passenger flow forecasting in urban rail transit, we developed a collaborative forecasting mechanism using ridge regression as the meta learner. During ridge regression, station- and line-level passenger flow forecasting is integrated within a unified optimization framework. At the meta-learning stage, it establishes relationships between the station- and line-level passenger flows to achieve joint forecasting optimization at the two levels.
3. Problem Description
Passenger flow data from rail transit stations are recorded sequentially in continuous time intervals to form a time series. Therefore, passenger flow forecasting of urban rail transit is a time-series forecasting problem, where historical data and explanatory variables are used to determine future values. The input of the forecasting model is the historical passenger flow sequences and the exogenous factors influencing the passenger flow, and its output is the future passenger flow.
The set of rail transit stations is represented by
, where
is the total number of stations,
represents the actual passenger flow of urban rail transit station
at time
, and
denotes its forecasting value. The objective of passenger flow forecasting for urban rail transit station
is to build a function
, as shown in Equation (1), that minimizes the forecasting error
across all training samples.
where
is the length of the historical data sequence and
denotes the external feature vector influencing the passenger flow of urban rail transit station
at time
.
Let
denote the line-level passenger flow at time
, and let
be the forecasting value. The function
that forecasts the line-level passenger flow at time
can be expressed in Equation (2), where
.
Passenger flows within urban rail transit have the following complex characteristics that make forecasting difficult:
(1) Passenger flows exhibit complex temporal patterns and strong non-stationarity, including long-term trends, multiple seasonalities and irregular fluctuations. These components are nonlinear and time-varying and influence each other. A single forecasting model cannot sufficiently address this complexity and its forecasting accuracy is hence reduced.
(2) Passenger flows are influenced by the multi-source spatio-temporal factors that shape their spatial distributions and temporal evolutions. Such factors include calendar effects (e.g., weekdays, weekends and holidays), station interactions (e.g., passenger transfers), and station functions (e.g., transit hub or commercial center). These factors interact with each other in a complex way, which makes the feature engineering of passenger flow patterns and designing forecasting models difficult.
(3) Multi-level forecasting has consistency issues. The operations, management, and planning teams within urban rail transit systems require support from both station- and line-level passenger flow forecasts. However, existing forecasting methods are only suitable for only one or the other and cannot deal with simultaneous forecasting at both levels.
4. Model Construction
To address the above difficulties, we propose a model that integrates multi-source spatio-temporal features and features optimized ensemble learning. The model sequentially performs spatio-temporal feature engineering, dynamically weighted ensemble forecasting, and meta-learning-based fusion to obtain station- and line-level passenger flow forecasts.
Multi-source spatio-temporal feature engineering is the core of the proposed model. It can perform forecasting using multiple passenger flow features to enhance the accuracy of the results. This feature engineering employs temporal feature extraction to characterize the temporal dependencies and periodic regularities in the passenger flows in urban rail transit. Empirical observations show that these passenger flows exhibit evident cyclical fluctuations, including distinct morning and evening peaks on weekdays and significant differences in the passenger flows between weekdays and weekends. Additionally, the passenger flow at a given time is strongly correlated with that observed at temporally adjacent intervals [
43]. All of these regularities influence passenger flow forecasting. To formulate these regularities in forecasts, the proposed model features the following two principal temporal features:
(1) Temporal periodic features that encode recurring cyclical patterns.
(2) Lag features that quantify the influence of the temporally adjacent historical observations on the target value.
These spatial features were constructed to characterize the complex interaction patterns among urban rail transit stations. The passenger flow at a given station is influenced by multiple factors (e.g., geographical location and station functions), and the resulting spatial effects on forecasting can be mainly classified into the following two categories:
(1) Propagation and diffusion of passenger flows between adjacent or spatially proximate urban rail transit stations.
(2) Coordinated variations that might occur between geographically distant urban rail transit stations due to similarities in functions and structural or operational characteristics [
44].
To address the spatial effects, we propose the following three spatial features:
(1) Graph adjacency features to encode the urban rail network topology with an adjacency matrix derived from the physical connections among urban rail transit stations and lines [
43]. This adjacency matrix is a binary matrix, where 1 indicates a direct connection between two stations and 0 indicates no direct connection. Accordingly, we can obtain a transfer matrix where each element represents the transfer probability of passenger flows between stations.
(2) Graph adjacency features alone cannot capture all spatial correlations, since non-contiguous stations might appear with strong interdependencies due to large events, key transfer hubs, or recurrent commuting patterns [
45]. Attention-based features are adopted in the proposed model to deal with this issue. Through such features, the query, key, and value are derived from the passenger flow. The key and value are set as identical to represent the same spatio-temporal state. The attention-based features are designed to infer and quantify the dynamic correlations between each pair of urban rail transit stations at a given temporal scale to overcome the rigidity of the static graph structures.
(3) Passenger flow evolution strongly depends on the function of a station [
46,
47]. To represent this, stations are grouped by their passenger flow profiles into clusters that reflect similar functions and passenger flow characteristics. Each station is then assigned a cluster label, which allows the forecasting model to distinguish station types and adapt to the flow dynamics.
The overall multi-source spatio-temporal feature engineering process is illustrated in
Figure 1.
We further established an ensemble learning model to overcome the single model’s lack of ability to jointly forecast noisy, nonlinear station- and line-level passenger flows that are influenced by multiple factors [
48]. We used ExtraTrees and LightGBM as the base models to build an ensemble learning model, since the former can avoid data overfitting and handle data uncertainty [
49] and the latter can model complex nonlinear relationships and feature interactions [
50]. The optimized ensemble learning outputs the forecasting results via the following procedure.
First, the station-level passenger flow for each station is initially forecasted by the two base models using the spatio-temporal features as inputs. We then employ the PSO algorithm to search for the optimal fusion weights of ExtraTrees and LightGBM to obtain the final forecasts, in which the robustness and accuracy are balanced.
Then, we sum the results of all stations in a line to obtain the aggregated station-level passenger flow forecasts. This process provides a bottom-up estimation of a line’s entire passenger demand while preserving station-level heterogeneity and enabling a line-wide analysis.
The initial line-level passenger flow can be forecasted by the ensemble learning model using the aggregated station-level passenger flows as training data, which enables line-wide temporal patterns, inter-station dynamics, and holistic line-level passenger flow trends to be captured.
Finally, we introduce a meta-learning fusion step to achieve an accurate final forecast. Ridge regression exhibits the advantages of reducing overfitting and handling multicollinearity among input features [
51]. Therefore, we employed it as the meta-learner in this step to forecast the final line-level passenger flow, taking the initial line-level and aggregated station-level passenger flow forecasts as the inputs.
The overall forecasting process of the ensemble learning model is shown in
Figure 2.
4.1. Temporal Features
Temporal features can either be discrete or continuous. The former are derived directly from timestamps, and include —the feature that captures weekly patterns—and —the feature that reflects daily variations.
However, discrete features cannot represent cyclical variations. Therefore, we used the continuous periodic features to map the actual time into a cyclic space using sine and cosine functions. This ensures that the same times on different days have the same representation, helping the model recognize and leverage repeated patterns across different time points. In this encoding, times near midnight remain close in the cyclic space so the discontinuity of discrete hour or minute indicators is avoided. By mapping time to a unit circle with two continuous features, the model captures the periodic patterns of daily passenger flows. Specifically, time
is first converted into a real number (e.g., 3:30 p.m. is converted into 930), and then transformed by Equation (3) into sine and cosine values:
where
and
are the sine and cosine representations of time
, respectively, and
is the total minutes in a day (i.e., 1440).
Lagged features identify short-term temporal dependences in passenger flows. Lag feature vector
for station
at time
was employed to capture the recent passenger flow, and is formulated in Equation (4):
where
denotes the lag length. Missing values at the start of
are padded with zeros. Finally, all temporal features of station
at time
are concatenated into a complete temporal feature vector
, as shown in Equation (5):
4.2. Spatial Features
4.2.1. Spatio-Temporal Attention Features
To capture spatio-temporal dependencies, we dynamically weighted the historical passenger flows across all the stations using an attention mechanism. This weight indicates the relevance of historical passenger flows to current passenger flow, which can be further used by the model. Unlike moving averages using fixed weights, the attention mechanism assigns weights based on the above relevance, and therefore the weights can be allocated more reasonably.
The historical passenger flow matrix at time
is represented by Equation (6):
where
is a row vector of the passenger flows for all stations at time
,
represents the length of the attention window, and
.
We took the last row of
as the query vector
at time
, i.e.,
, and used
as both the key matrix
and value matrix
at time
, i.e.,
,
and
. The attention score, that is, the scaled dot product, is computed by Equation (7):
where
is the dimension of the key matrix.
To ensure numerical stability, we subtracted the maximum value of the attention score, as shown in Equation (8):
Attention weight
represents the importance of historical passenger flows within the window of the current forecast, and is obtained by Equation (9):
where
denotes the index of the historical passenger flow matrix at time
. The final attention output is a weighted sum of the value matrix, as shown in Equation (10):
where
, in which
represents the adaptively weighted value of passenger flow for station
at time
.
4.2.2. Graph Adjacency Features
The graph adjacency features are determined by the connectivity of the urban rail transit network, and can thus simulate passenger flow propagation between the stations in the network. Passenger flow moves between two adjacent stations or lines, which has an influence on both stations’ individual flows. Modeling these relationships is crucial for passenger flow forecasting.
Spatial connectivity is encoded in an
adjacency matrix
, whose elements are given by Equation (11):
where
represents a direct connection allowing passenger flow movement from
to
. To represent flow, the transfer probability,
, is row-normalized to obtain an
transfer matrix. The transfer matrix
represents the transfer probabilities of passengers moving from station
to station
, and is calculated by Equation (12):
After obtaining the one-step transfer matrix
, we further leverage this matrix to capture the multi-step propagation of passenger flows. Let
denote the matrix raised to power
, and its elements
represent the probability that the passenger flow moves from station
to station
in
steps. For a given
, spatial propagation feature vector
at time
is given by Equation (13):
where
is the passenger flow vector of all stations at time
. The
-th power spatial feature of station
at time
is determined by Equation (14):
where
denotes the index of the elements in
, and
represents the passenger flow arriving at station
at time
after
steps.
4.2.3. Station Clustering Features
Urban rail transit stations are clustered based on passenger flows. This allows the model to use the shared characteristics of stations within the same cluster when forecasting passenger flows, which improves its generalization.
When constructing the spatial clustering features, the passenger flow of rail transit stations should first be standardized. The K-means algorithm should then be applied to group stations with similar passenger flow characteristics: after initializing cluster centers, each station is assigned to the nearest centroid based on the Euclidean distance, and centroids are recalculated until convergence. Upon completion, each station is assigned a final cluster label.
Based on the clustering result, two features are created for each station: (1) a cluster historical aggregate passenger flow , that is, the total passenger flow within the cluster label to which station belongs, and (2) a one-hot encoding vector which identifies the cluster membership of station .
The final output of this stage is a complete spatial feature vector for station
at time
, formed by concatenating attention mechanism output
, graph feature
, and the two cluster features:
The procedure of spatial clustering feature construction is shown in
Figure 3.
4.3. Ensemble Learning
We developed a two-stage weighted ensemble framework to integrate the multi-source spatio-temporal features obtained in
Section 4.1 and
Section 4.2. In this way, reliable and stable passenger flow forecasts were obtained. In Stage 1, two base learners, ExtraTrees and LightGBM, were combined by using dynamically optimized weights to generate the initial station-level forecast. The input to each model was the spatio-temporal feature vector
of station
at time
, which is presented in Equation (16). In Stage 2, a meta-learner, i.e., ridge regression, was employed to integrate the initial line-level forecast with the aggregated station-level forecasts to obtain a final line-level forecast.
4.3.1. Base Model I: ExtraTrees
ExtraTrees is a highly randomized tree ensemble. It builds multiple decision trees by randomly selecting feature subsets and split points to carry out cooperative forecasting. Each tree is built with different feature and split choices. This randomness reduces variance and improves generalization. The structure of ExtraTrees is shown in
Figure 4.
The final station-level forecasting result of station
at time
is the average output of all trees’ outputs, as shown in Equation (17):
4.3.2. Base Model II: LightGBM
LightGBM is an efficient gradient-boosting framework. It grows trees sequentially, and each new tree is trained to correct the residual errors of the current ensemble. LightGBM uses histogram-based algorithms to achieve a fast training speed and a leaf-wise growth strategy to achieve high accuracy. Therefore, it is suitable for processing large amounts of data. The structure of LightGBM is shown in
Figure 5.
The station-level passenger flow forecast for station
at time
is given by Equation (18):
4.3.3. Dynamic Weight Optimization via the PSO Algorithm
The weighted combination of the two base models produces the final station-level forecast. The optimal weights are determined by minimizing the forecasting error of the validation set by using the PSO algorithm.
The validation set is denoted as
. The combined forecasting result
using the weight vector
is presented in Equation (19):
where
and
are the weights assigned to the forecasting results given by ExtraTrees and LightGBM, respectively. The objective of Equation (20) is to find the optimal weight vector
that minimizes the Root Mean Square Error (RMSE) of the validation set:
The PSO algorithm is employed to solve this optimization problem. A swarm of
particles is initialized. At iteration
, each particle
has a position vector
, which represents a candidate weight vector and a velocity vector
. Particles update their states based on their personal best position
and the swarm’s global best position
, which are updated at each iteration
, as indicated by Equations (21) and (22):
where
is the inertia weight,
and
are acceleration coefficients, and
and
are random numbers that are uniformly distributed between 0 and 1. Upon convergence, the final global best position
is set as
. Then,
is put into Equation (23) to normalize the weights to ensure the sum of their values is 1 in order to ensure that the relative contribution of each base learner is fully represented in the final station-level passenger flow forecast.
where
is a small enough constant for numerical stability.
and
represent the final normalized combination weights.
Consequently, the final station-level passenger flow forecast
of station
at time
is:
4.3.4. Meta-Learning for Multi-Level Forecasting Reconciliation
Ridge regression was used as the meta-learner to obtain the final line-level forecast. We first aggregated all station-level forecasts to obtain the aggregated station-level forecast based on Equation (25):
The proposed ensemble learning model (
Section 4.3.1,
Section 4.3.2 and
Section 4.3.3) was adopted again to produce the initial line-level forecast
of line
using the same multi-source spatio-temporal features.
The meta-learner was trained by dataset , where is the total number of time steps in the dataset and is the actual line-level passenger flow of line .
The model learns a coefficient vector
according to Equation (26). This vector minimizes the loss of regularized forecast error:
where
is the intercept,
is the weight for the aggregated station-level forecast,
is the weight of the initial line-level passenger flow forecast, and
is the regularization strength.
penalizes large coefficients to prevent overfitting and is equal to
.
After training using a meta-learner, the optimal coefficients
,
and
are obtained. The final forecasted line-level passenger flow forecast
for line
at time
is shown in Equation (27):
5. Model Validation
5.1. Data
We employed passenger flow data obtained from smart card transactions in Hangzhou’s urban rail transit system to verify the proposed forecasting model. The data are available at Alibaba Cloud (
https://tianchi.aliyun.com/dataset/21904, accessed on 10 January 2026). The original dataset contains 70 million individual transaction records from 1 January 2019 to 26 January 2019. Each record details the transaction time; station ID, which ranges from 0 to 80; and line ID (A, B or C). Specifically, Line A covers stations 67 to 80, Line B covers stations 0 to 33 and Line C includes the remaining stations. Additionally, the dataset contains an adjacency matrix that encodes the topological relationships of the stations based on the tracks connecting them.
We aggregated the raw transaction data during operational hours (from 5:30 to 23:30) into 15-min intervals to construct a passenger flow time-series for each station. The actual line-level passenger flow was calculated by aggregating the actual passenger flows of all the stations belonging to the same line for each 15-min interval.
After constructing the passenger flow time-series, we chronologically split the dataset into training and testing sets with a ratio of 8:2. The first 80% was used for model training and parameter tuning, and the rest for model evaluation. All variables were standardized using training set statistics, and each station’s validation data was standardized using the mean and standard deviation of its training data to avoid information leakage.
After the above processing, the station- and line-level passenger flows, recorded in 15-minute intervals, were considered as the forecasting target. The passenger flow data for Lines A, B and C are shown in
Figure 6.
As shown in
Figure 6, the passenger flows on Lines A, B and C all exhibited cyclical fluctuations that are associated with commuting patterns on weekdays and weekends. Such fluctuations remained consistent across the training and validation periods.
Table 1 presents the statistical characteristics of urban rail transit passenger flows in Lines A, B and C using the following statistical measures: “count”, which denotes the total number of passenger flows; “mean”, which represents the average; “std”, the standard deviation, “min”, which is the minimum passenger flow; “Q1”, “Q2”, and “Q3”, which indicate passenger flows in the first, second and third quartiles, respectively; and “max”, the maximum passenger flow. All values are expressed as passenger flow (persons).
As shown in
Table 1, Line B had the highest average passenger flow and the largest fluctuations, which means that it has a highly variable passenger demand. Line C had a moderate average passenger flow and variability, while Line A had the minimum average passenger flow and the smallest fluctuations. All three lines experienced periods of considerably low passenger flow and severe peak-hour congestion.
Overall, the passenger flow characteristics of the three lines are different to each other. An effective forecasting model should be able to adapt to the features of different lines; therefore, studying these three lines can demonstrate the adaptivity and generalization of our forecasting model.
5.2. Lag Length Determination
To determine an appropriate lag length for capturing short-term dependencies, we examined the temporal autocorrelation structure of the passenger flow time-series. The autocorrelation function (ACF) and partial autocorrelation function (PACF) were computed and plotted for each of the three lines. These plots provide insight into the extent to which current passenger flows are correlated with past flows, thereby guiding the selection of the lag length used in Equation (4).
The ACF and PACF of passenger flows for Lines A, B, and C are presented in
Figure 7. The ACF revealed a clear diurnal pattern: there was a rapid decay within the first five lags, negative troughs at lags 6–28, and two positive peaks at lags 32–40 and 64–72, with the latter exceeding 0.6. The PACF exhibited extremely high values at lag 1, sharply dropped at lag 2, and became statistically insignificant after lag 5 for Lines A and B, while Line C showed slight significance at lag 3.
These results demonstrate two key characteristics of the data: strong short-term dependence, captured within five lags, and pronounced diurnal periodicity. Accordingly, setting in Equation (4) captures all significant short-term correlations, while sine–cosine encoding efficiently represents the periodic patterns. This combined design ensures comprehensive feature expression while avoiding the computational burden and feature redundancy that would result from including dozens of lag terms, thereby reducing model size and maintaining its efficiency.
5.3. Cluster Analysis
To construct spatial features, we employed a K-means cluster algorithm to classify the urban rail transit passenger flow into four categories. The resulting four clusters represent distinct passenger flow patterns that implicitly reflect different station functions, such as residential areas, commercial districts, transportation hubs, or mixed-use zones [
40]. The proposed model can accordingly forecast the passenger flows by considering clusters containing unique station characteristics, which helps to improve the forecasting accuracy.
To better visualize the station clusters of the three urban rail transit lines, we reduced the original high-dimensional passenger flow features using Principal Component Analysis (PCA). All passenger flow features were projected into the first and second principal components to capture the greatest variance in the original flow data. The clustering results of the three lines are shown in
Figure 8.
As illustrated in
Figure 8, stations on Line A were dispersed across the plot, which reflects the substantial variation in their passenger flows. For Line B, the stations were spread widely along the first principal component, with Cluster 0 in the positive region and Cluster 2 in the negative region. This indicates a large overall difference in passenger volumes among the stations on Line B. In contrast, the stations on Line C were closely grouped together, with Clusters 0 and 3 overlapping near the origin, which suggests that passenger flows are more uniform across Line C.
Table 2 lists the station clusters for Lines A, B and C.
The cluster labels (i.e., 0, 1, 2 and 3) of these urban rail stations were used as input (i.e., an explanatory variable) to forecast the station- and line-level passenger flows.
5.4. Comparative Experiment
5.4.1. Selected Benchmark Models for Comparative Experiments
To evaluate the proposed model, we compared it with four advanced deep learning and eight traditional machine learning benchmark models. The deep learning benchmark models, able to capture temporal dependencies, are as follows:
(1) Gated recurrent unit (GRU) with gating and lower complexity;
(2) LSTM network, which learns long-range dependencies via cell states and gates;
(3) Transformer, which uses self-attention to model the global temporal relationships;
(4) Temporal convolutional network (TCN), which applies causal and dilated convolutions to obtain a wide receptive field with clear temporal order and parallel computation.
The machine learning benchmark models are tree- and kernel-based ensemble models that exhibit strong performance on structured data. All machine learning models were trained using default parameters and are listed below:
(1) Adaptive Boosting (AdaBoost), which sequentially focuses on mispredicted samples;
(2) Categorical Boosting (CatBoost), which handles categorical features and mitigates overfitting by ordered boosting;
(3) ExtraTrees, which adds randomness to tree construction to reduce variance;
(4) Gradient Boosting Regression (GBR), which builds an additive model to minimize a differentiable loss;
(5) LightGBM, which uses efficient histogram-based and leaf-wise gradient boosting;
(6) Random forest (RF), which aggregates multiple trees trained by bootstrapped samples and random feature subsets;
(7) Support Vector Regression (SVR), which fits data within a tolerance margin via kernels in high-dimensional space;
(8) Extreme Gradient Boosting (XGBoost), which performs regularized gradient tree boosting with high speed and accuracy.
The parameter settings of all the benchmark models are listed in
Table 3; the parameter settings of the model proposed in this study are given in
Table 4.
As shown in
Table 4, the key spatio-temporal parameters were selected empirically from the data. For graph propagation, the transfer matrix order was set to 2, capturing both direct neighbors and stations two steps away. The attention window size was set to 4, corresponding to a one-hour history with 15-minute data, which captures short-term dependencies and recent fluctuations while remaining efficient. These settings allow for underlying patterns to be represented by features without overfitting or introducing unnecessary complexity.
5.4.2. Comparative Results
In this study, we evaluated model performance using the Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and the Coefficient of Determination (R2). The RMSE measures the deviation between the forecasted results and actual values and is sensitive to large errors, while the easily interpretable MAE represents the average absolute forecasting error and the MAPE expresses the forecasting accuracy as a percentage, which allows for a comparison across different data scales. Finally, R2 is the proportion of variance explained by the model, where a value closer to 1.0 indicates a better fit.
(1) Station-level passenger flow forecasting results
The forecasting errors (i.e., RMSE, MAE, MAPE and R
2) of the proposed model and the selected benchmark models for station-level passenger flows are presented in
Table 5. It should be noted that
Table 5 presents the average forecasting errors of all stations on a line.
According to
Table 5, the proposed model achieved the minimum RSME, MAE and MAPE in all cases. It also yielded the maximum R
2 for Lines A and C and a near-maximum R
2 for Line B. Therefore, we can conclude that the proposed model exhibited the best forecasting performance on all three lines compared to the benchmark models.
It should also be noted that there were much larger forecasting errors for Line B than Lines A and C across all models, since this line is a major corridor with higher passenger flow (see
Table 1) and contains large transfer hubs. However, the proposed model still achieved the minimum RMSE, MAE and MAPE, which demonstrates that it is able to deal with considerable passenger flows through urban rail transit stations and thus shows high adaptivity.
Figure 9 comprehensively illustrates the capability of the proposed model in capturing the change in passenger flows at all studied stations on Lines A, B, and C.
For the stations on Line A, passenger flows showed clear daily rises and falls with distinct peak hours, which reflect typical commuter cycles. Stations 67, 71, 75–78 particularly exhibited relatively high passenger flows during some peak hours, while the passenger flows of Stations 70 and 74 remained lower and steadier. According to
Figure 9a, the proposed method can effectively capture the temporal correlations and daily cycles in the passenger flow data for all stations on Line A.
For the stations on Line B, passenger flows varied greatly. These stations include not only high-volume transfer hubs with severe crowding (e.g., Station 15) but also small-scale stations with lower passenger demand (e.g., Station 31). The spatial links between the stations in Line B are thus complicated. However, based on
Figure 9b, the proposed model still successfully learnt the above characteristics in passenger flows for all stations on Line B.
For the stations on Line C, passenger flows were more regular and stable compared to Lines A and B. Daily commute trends were consistent across stations on Line C, which exhibited similar peak hours and less variation. Moreover, Line C had more uniform passenger flows, which makes station-level forecasting easier. As indicated by
Figure 9c, the proposed model can accurately match the above regularity, stability and consistency.
Overall, the proposed model can accurately capture the variations in station-level passenger flows in both regular and peak periods, and exhibits accurate forecasting for urban rail transit stations with different passenger flow characteristics.
(2) Line-level passenger flow forecasting results
The line-level passenger flow forecasts for the benchmark models were obtained by summing the forecasted results (from the same model) of all stations on the line. In contrast, the proposed model used meta-learning to optimize the forecasts of both station- and line-level passenger flows. The line-level passenger flow forecasting errors of the proposed model and benchmark models are presented in
Table 6.
As shown in
Table 6, the proposed model achieved the best performance in forecasting line-level passenger flows compared to the benchmark models, yielding the minimum RMSE, MAE and MAPE and maximum R
2 for all three lines. Thus, the combination of multi-source spatio-temporal features and optimized ensemble learning is effective in forecasting the line-level passenger flows of urban rail transit systems.
Figure 10 visually illustrates the performance of the proposed model in line-level passenger flow forecasting.
The passenger flow of Line A was the lowest among the three lines, exhibiting moderate peak hours and clear drops in passenger flow on weekends. Additionally, passenger volumes differed greatly between weekdays and weekends, and weekly changes were gentler during off-peak hours. As illustrated in
Figure 10a, the proposed model is a suitable choice for urban rail transit lines with the above characteristics.
For Line B, passenger flow was the highest, and there was severe overcrowding during the morning and evening peak hours and large decreases during the off-peak hours. Its passenger flows and intensity were considerably greater than those of Line A. However, according to
Figure 10b, the proposed model could still handle this situation well.
For Line C, passenger flow and variation were intermediate compared to Lines A and B.
Figure 10c demonstrates that the proposed model still effectively handled these features to achieve an accurate forecast.
Above all, the close alignment between the actual and forecasted curves for all three lines demonstrates the ability of the proposed model to capture the overall temporal dynamics of and fluctuations in line-level passenger flows for urban rail transit lines with varying passenger flow characteristics.
The results indicate that meta-learning based on ridge regression can effectively optimize and balance station- and line-level passenger flow forecasts to achieve consistency between them. The proposed model is effective in forecasting these flows across both levels, and exhibits high adaptivity and generalizability to handle stations and lines with different passenger flow characteristics.
5.5. Performance Analysis on Station Clustering
To further investigate the forecasting performance across different station types, we analyzed the station clustering results on Line B. As described in
Section 5.2, the stations on Line B were grouped into four clusters based on their passenger flow patterns. For each cluster, we selected the station with the lowest RMSE as a representative and plotted its actual versus forecasted passenger flows over the test period, as shown in
Figure 11.
As shown in
Figure 11, cluster 0 contained 14 stations typical of residential or peripheral areas, with sharp morning peaks of above 1600 passengers, near-zero overnight flows, and rapid evening declines. The model closely tracked the steep morning rise and evening drop, matching the commute-driven demand profile. Even when volumes were near zero in early morning hours, forecasts remained tightly aligned with the observations, indicating that the model effectively learns station-specific diurnal rhythms.
Cluster 1 included 15 stations functioning as high-volume central business district nodes or major transfer hubs. Their passenger flow profiles featured sustained loads ranging from 3000 to 4500 passengers throughout the day, with flattened peak structures that lacked the sharp spikes seen in residential stations. The model achieved strong alignment during both morning and evening peak periods, and maintained reasonable accuracy during off-peak hours. This consistent performance reflects the stable, high-density demand characteristic of commercial centers, where passenger activity remains elevated across a broad daytime window rather than concentrating into narrow peaks.
Cluster 2, the largest group with 31 stations, displayed moderate-volume mixed-use patterns. These stations exhibited irregular peak structures and higher intraday volatility compared to the other clusters, indicative of diverse trip purposes, including commuting, shopping, and leisure activities. Despite this complexity, the model successfully followed the overarching trend and captured the principal fluctuations throughout each day. The forecasts adapted well to the variable demand patterns, underscoring the model’s robustness when faced with heterogeneous station functions.
Cluster 3 contained nine stations with evening-dominated patterns typical of entertainment districts or commercial areas oriented toward nighttime activity. Passenger flow gradually ascended during the afternoon and accelerated through the early evening, culminating in pronounced late-evening peaks exceeding 3000 passengers. The model accurately reproduced this distinctive temporal shift, tracking the evening surge with precision and maintaining fidelity throughout the extended active period.
In summary, the model fit morning commute clusters tightly and performed best during the peak period for evening entertainment clusters. The off-peak daytime accuracy stayed high in business districts, and the model exhibited low deep off-peak absolute errors, aligned with reduced demand. These results have direct operational implications: reliable peak forecasts can be used in high-volume hubs for dynamic headway adjustments; in residential stations to optimize early-morning deployment; in mixed-use areas to integrate real-time data; and in evening-centric stations to justify increased late-night train frequencies.
6. Conclusions
This study addressed the problem of forecasting passenger flow in urban rail transit at the station and line levels, and a novel forecasting model was proposed that integrates multi-source spatio-temporal features and optimized ensemble learning. The main contributions of this study are as follows:
(1) A comprehensive multi-source spatio-temporal feature engineering framework was proposed to effectively capture the complex and nonlinear dynamics of station-level passenger flow. It systematically extracts and fuses different features, including temporal periodic features, lag features, and spatial features derived from graph adjacency; an attention mechanism; and station clustering.
(2) A dynamically weighted ensemble forecasting framework was developed to enhance robustness and accuracy. It employs ExtraTrees and LightGBM as base models and utilizes the PSO algorithm to adaptively determine their optimal combination weights based on optimizations of the validation performance.
(3) A meta-learning-based collaborative forecasting mechanism was introduced to optimize and balance station- and line-level forecasts. Ridge regression was used as the meta-learner and trained to process station-level forecasts in order to obtain line-level forecasts.
(4) A systematic validation experiment based on passenger flow data from three urban rail transit lines in Hangzhou was conducted to verify the advantages of the proposed model. Through comparative experiments, it was verified that the proposed model yields higher accuracy and robustness and is more applicable to real situations.
The results of this study show that the proposed model can provide solid decision support for operations, management, and planning departments across stations and the entire network, including support for real-time crowd management, dynamic train scheduling, and efficient resource allocation.
However, the proposed model still has the following limitations that should be addressed in future research.
(1) The proposed model relies on passenger flow data extracted from smart card transactions and does not incorporate external geospatial information. Due to the nature of the public dataset, the specific physical alignments of the three lines were not explicitly defined, preventing us from inferring the precise geographical locations of stations and their surrounding land use characteristics. Consequently, important exogenous factors such as points of interest, employment density, and residential distribution—which significantly influence passenger demand—could not be integrated into the feature engineering stage.
(2) The proposed model groups stations using clustering based solely on passenger flow time-series, without considering functional classifications such as transfer hub, commercial center, or residential area. The lack of ground-truth station function labels, again attributable to the absence of external data sources, limits our ability to validate whether the identified clusters truly reflect operational roles within the network.
(3) The proposed model was validated using a public dataset containing data across a single month (January 2019), which may not fully capture seasonal variations, holiday effects, or long-term ridership trends. No additional data from different seasons or years were available to assess the model’s robustness under varying temporal conditions.
To address the above limitations, future research will be conducted in the following directions.
(1) To address the lack of external geospatial information, we will integrate smart card data with multi-source geospatial datasets, including POIs, land use maps, and demographic statistics. Using spatial analysis methods such as geographically weighted regression or graph neural networks with external node features, we will quantify how urban forms and land use affect passenger flows, improving the model’s explanatory power and accuracy.
(2) In future work, we will refine station clustering by adding functional classification variables from external data, including POI data, transfer station attributes, and employment density. We will develop hybrid clustering methods that combine passenger flow time-series with these indicators to better represent stations’ operational roles and enhance the interpretability of cluster-based features.
(3) We will strengthen the model’s robustness by validating it on longer-term datasets covering multiple seasons and years, enabling the analysis of seasonal fluctuations, holiday effects, and long-term ridership trends. We will also incorporate time-varying external factors, such as weather and special events, into the feature engineering stage so that the model can capture more complex temporal dynamics and be more useful in practice.