1. Introduction
With the acceleration of global urbanization, rapid urban population growth, and continuous urban spatial expansion, urban transportation systems are facing unprecedented challenges. In this context, urban rail transit has emerged as a key solution to urban transportation problems due to its efficiency, environmental friendliness, and high capacity. As critical nodes in urban transportation networks, rail transit stations not only provide convenient travel services but also profoundly influence urban spatial structure and residents’ travel patterns [
1,
2,
3]. The differentiated functional characteristics of stations directly affect the spatiotemporal distribution of passenger flow, thereby altering the overall efficiency of urban transportation networks. Therefore, scientifically rational classification of rail transit stations and analysis of their passenger flow spatiotemporal differentiation patterns hold significant theoretical value and practical importance for optimizing urban public transportation structures and enhancing service effectiveness.
Current research on the classification of urban rail transit (URT) stations is already extensive both domestically and internationally. In terms of delineating station influence areas, public transportation studies commonly establish an 800 m distance as the conventional threshold for pedestrian access, widely applied to define the service scope of urban rail transit stations [
4,
5,
6,
7]. For instance, rail transit stations were categorized into five types based on land-use functions within an 800 m radius [
8]. However, a fixed-radius approach is not universally applicable across all urban contexts, particularly in high-density areas where it often leads to overlapping influence zones between adjacent stations, compromising classification accuracy. To address this issue, some scholars have introduced Voronoi diagrams as an alternative to the traditional 800 m buffer, aiming to resolve overlaps in station influence areas [
9,
10]. While this method effectively eliminates zonal overlaps, it overlooks disparities in area size, resulting in inadequate characterization of station heterogeneity and raising concerns regarding evaluation fairness and comparability.
Secondly, recognizing the complex interrelationships among stations, built environment, and passenger flows, the scientific classification of urban rail transit stations to reveal their intrinsic functional heterogeneity has emerged as a critical research direction, prompting extensive exploration [
11,
12,
13,
14,
15]. For instance, Duan et al. classified stations by focusing on land use and land-use uniformity, proposing corresponding optimization strategies for surrounding land-use patterns [
16]. Papa E et al. conducted a typological study of station influence areas using cluster analysis, suggesting public transport priority strategies to mitigate the negative impacts of cars [
17]. They developed a method for identifying station area types and validated it in Naples, demonstrating how coordinated land-use and transportation planning can enhance rail station efficiency. Gan et al. investigated the relationship between urban built environment characteristics including land use, transportation facilities, population density and subway passenger flow, with a focus on analyzing the influencing factors and their interactions among different subway stations [
18]. Huang et al. explored the relationship between the built environment characteristics of Beijing Metro stations and transit-oriented development (TOD), using regression analysis to reveal how environmental factors influence subway ridership, providing empirical evidence for urban transport planning [
19]. Furthermore, conventional classification methods often rely on small-scale survey data, particularly those considering only a single time period, which inadequately captures temporal variations [
20]. In response, several scholars have developed clustering approaches for multiple time-series data [
21,
22,
23]. However, these methods overlook correlations among multivariate time-series variables and may group morphologically divergent time-series sequences into the same cluster. To address these limitations, Zhang et al. modeled passenger flow as time-series curves and introduced a two-stage multivariate time-series clustering method, focusing exclusively on curve characteristics to classify urban rail transit stations [
24]. However, most existing studies rely on static indicators or focus solely on built environment features for station clustering. Even when time-series data is considered, the analysis is often limited to a single dimension, lacking a multi-dimensional spatiotemporal approach to station classification that accounts for differentiated spatiotemporal characteristics.
With the advancement of machine learning technologies, clustering algorithms have been increasingly applied in the transportation field, among which methods such as K-Means Clustering Algorithm (K-Means) and Density-Based Spatial Clustering of Applications with Noise (DBSCAN) are widely adopted for station classification. The K-Means algorithm enhances intra-cluster similarity by minimizing within-cluster variance. Recognized for its straightforward iterative process and efficiency in rapidly identifying stations with similar characteristics, it has become a conventional method in rail transit station classification [
25,
26,
27,
28]. In contrast, DBSCAN, as a density-based clustering algorithm, excels at identifying clusters of varying shapes and sizes based on data point density. Its strengths lie in handling noise and outliers, and unlike traditional K-Means, DBSCAN can discern non-spherical station clusters in transportation networks [
29]. Furthermore, analytical methods for studying passenger flow influence mechanisms at stations are primarily categorized into global regression and local regression approaches. The most commonly used global regression method is Ordinary Least Squares (OLS) [
30]. Some scholars, considering the geospatial characteristics of influencing factors, have adopted the typical local regression model Geographically Weighted Regression (GWR) to address the limitation of fixed coefficients in OLS models [
31,
32,
33]. In recent years, scholars both domestically and internationally have begun to employ Multiscale Geographically Weighted Regression (MGWR) models to investigate the influencing factors of station passenger flow in greater depth. Compared to GWR models, MGWR models offer the flexibility to set varying bandwidths for different influencing factors [
34,
35,
36]. However, traditional clustering algorithms and single regression models fail to account for the spatial heterogeneity of actual stations and local variations in influencing factors, which not only compromises classification accuracy but also makes it difficult to distinguish the significant influencing factors for different station types.
Table 1 provides a detailed summary of the key algorithms and findings discussed above.
In summary, existing research on rail transit station classification and passenger flow characteristics still has several limitations: On the one hand, the delineation of influence areas predominantly relies on fixed-radius buffers or Voronoi diagrams, which often leads to overlapping influence zones between adjacent stations or introduces fairness issues in subsequent clustering due to disparities in zonal areas. These limitations weaken the characterization of station heterogeneity. Furthermore, in terms of clustering indicator selection, traditional methods predominantly rely on static indicators or single temporal-spatial dimensions for classification. These approaches fail to adequately account for the dynamic variations in passenger flow patterns or the spatiotemporal coupling relationships between passenger flow characteristics and station access effects. On the other hand, conventional clustering algorithms and single regression models struggle to effectively capture the spatial heterogeneity among stations and local variations in influencing factors. As a result, the classification outcomes exhibit low consistency with actual station functionalities, and the models demonstrate limited explanatory power. To address the aforementioned issues, this paper proposes a non-overlapping zoning algorithm to delineate station influence areas. By integrating multidimensional indicators—including dynamic passenger flows, resident attributes, connection characteristics, and other relevant factors—a station clustering model based on the enhanced PAM algorithm is developed. Subsequently, OLS, GWR, and MGWR models are employed to analyze the spatiotemporal distribution patterns of passenger flows across different station categories. The findings provide a methodological reference for the refined management of urban rail transit stations.
To make the research objective more explicit, the overall objective of this study is to develop a replicable and interpretable analytical framework for the functional classification of rail transit stations and for examining the spatiotemporal heterogeneity of passenger-flow determinants. Specifically, the study first defines non-overlapping station influence areas as spatial analytical units and clarifies four groups of variables: dynamic passenger flow, resident attributes, connection characteristics, and spatial distribution. It then identifies station functional types through feature standardization, Principal Component Analysis (PCA), and cosine-distance-based PAM clustering, and finally compares the global and local effects of passenger-flow determinants across station types using OLS, GWR, and MGWR models. This design ensures that the classification results are supported not only by clustering outputs but also by objectively defined variables, statistical tests, and spatial interpretation.
The theoretical contribution of this study is that it extends station classification from a conventional “fixed spatial unit-static indicator clustering” approach to an integrated framework of “non-overlapping spatial units, multi-scale feature fusion, and spatial heterogeneity interpretation.” Previous studies have often focused on land-use indicators or passenger-flow curves alone, while paying less attention to the simultaneous problems of overlapping station catchment areas, redundancy among classification variables, and scale-varying passenger-flow determinants across station types. By combining non-overlapping influence areas, an enhanced PAM algorithm, and MGWR analysis, the proposed framework provides both a functional typology of stations and an explanation of the spatial mechanisms behind passenger-flow variation.
From a practical perspective, functional classification is needed in several planning and management contexts. For example, during new-line planning or renewal of existing lines, the classification can identify commuting-oriented stations, basic-service stations, leisure-vitality stations, and mixed-use stations. For daily operations, transport authorities can use the results to design weekday/weekend-specific service plans, bus feeder coordination, bike-sharing allocation, and passenger-flow management strategies. For investment prioritization and Transit-Oriented Development (TOD) policies, the classification can help determine where to improve local services and walking/cycling access, and where to enhance transfer capacity and public-space resilience.
The paper is structured as follows: Following this introduction,
Section 2 determines station influence areas,
Section 3 clarifies the variable system and multidimensional feature engineering used for station classification and passenger-flow analysis,
Section 4 presents the cosine-distance-based PAM clustering algorithm and the OLS, GWR, and MGWR models,
Section 5 conducts the Beijing case study and verifies class differentiation, and
Section 6 summarizes the conclusions, limitations, and methodological replicability.
6. Conclusions
Taking the Beijing urban rail transit system as a case study, this research develops an integrated station functional classification framework that combines non-overlapping station influence areas, enhanced PAM clustering, and spatial regression models. Methodologically, the study shows that non-overlapping influence areas reduce variable mixing between adjacent stations, multi-scale feature fusion captures functional differences more comprehensively, cosine-distance-based PAM clustering improves classification robustness, and MGWR reveals the spatial operating scales of different passenger-flow determinants. Empirically, Beijing rail transit stations are classified into four weekday types, A Peripheral Basic-Service Type, B Core Commuting-Aggregation Type, C Exurban Residential-Transit-Dependent Type, and D Multifunctional-Complex Type, and three weekend types, a Peripheral Living-Service Type, b Core Leisure-Vitality Type, and c Central Mixed-Use Type. Practically, the classification results can support differentiated service planning, feeder bus and bike-sharing integration, station-renewal investment prioritization, passenger-flow management, and TOD function allocation.
The findings of this study provide direct evidence to inform station planning, design, and management: (1) Differentiated strategies should be adopted according to station type. “Core Type” stations require ensuring connectivity within high-density road networks and improving public transport transfer efficiency to cope with tidal passenger flows. “Peripheral Type” stations should focus on compensating for the lack of lifestyle service facilities (POI) and enhancing the accessibility of public services. (2) The analysis of connection characteristics offers strategies for multi-modal transport integration and provides a precise decision-making basis for allocating interchange facility resources across different station categories. (3) The disparity in passenger flow drivers between weekdays and weekends indicates that Transit-Oriented Development (TOD) planning should enhance the integration of “commuting-leisure” functions. Core areas should increase leisure functions to boost weekend vitality, while residential areas should strengthen community services to promote all-day passenger flow balance. In addition, this study has several limitations and directions for future extension. First, the empirical analysis is based only on Beijing, and the applicability of the findings to cities with different urban scales, rail network structures, and travel behavior patterns requires further validation. Second, although this study integrates multi-scale features including dynamic passenger flow, resident attributes, connection characteristics, and spatial distribution, the representation of built-environment factors such as land-use mix, development intensity, population/employment density, and multimodal public transport coupling remains insufficient. Nevertheless, the proposed workflow, including non-overlapping station influence-area delineation, multidimensional feature construction, cosine-distance-based PAM clustering, and comparative OLS/GWR/MGWR analysis, is methodologically replicable. For cities with comparable AFC, POI, road network, resident attribute, and transfer-mode data, this framework can be applied to identify station functional types and examine the spatial heterogeneity of passenger-flow determinants. Future research will expand the analysis to multiple cities and richer data sources to further test the generalizability and robustness of the proposed method.
Based on these findings, station types can be used as basic units for refined rail-transit management. For B and D stations, planning and operation should focus on transfer efficiency, station-area circulation, and peak-flow organization. For C and a stations, priority should be given to feeder bus services, walking/cycling access, and local service facilities. For b and c stations, flexible weekend and holiday capacity, bike-sharing dispatch, and public-space management should be strengthened. This classification-diagnosis-governance chain provides a practical basis for coordination among transport authorities, rail operators, and TOD planners.