Next Article in Journal
Hierarchical Clustering and Schur Complement for Automatic Hyperspectral Band Selection
Previous Article in Journal
Surrogate Modeling and Optimization of a Dual-Band Circular Patch Antenna with a C-Shaped Slot Using MLP Neural Networks
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Context-Gated Graph Modelling for Traffic Flow Forecasting

by
Yuzhuo Zhang
1,*,
Jialin Liang
1,
Ziqiong Yuan
2,
Zanzan Dai
1 and
Yaozheng Kang
1
1
College of Vehicle and Traffic Engineering, Henan University of Science and Technology, Luoyang 471003, China
2
School of Transportation and Logistics Engineering, Wuhan University of Technology, Wuhan 430070, China
*
Author to whom correspondence should be addressed.
Modelling 2026, 7(4), 157; https://doi.org/10.3390/modelling7040157
Submission received: 29 June 2026 / Revised: 27 July 2026 / Accepted: 3 August 2026 / Published: 5 August 2026
(This article belongs to the Section Modelling in Artificial Intelligence)

Abstract

Traffic states evolve on irregular sensor graphs and vary with calendar context, yet the original ASTGCN does not explicitly model how the contribution of different graph receptive fields changes across traffic periods. This paper proposes CD-MRFG, a context-gated extension of ASTGCN that encodes hour-of-day, day-of-week and weekend information and uses the resulting representation to weight Chebyshev graph-convolution orders in each spatio-temporal block. Under a common 12-step forecasting protocol, CD-MRFG reduced the overall MAE and RMSE of the reproduced ASTGCN baseline from 18.66 and 31.05 to 16.98 and 28.59 on PEMS03, from 22.79 and 35.02 to 20.82 and 32.77 on PEMS04, and from 18.88 and 28.83 to 17.24 and 26.84 on PEMS08. Three-seed experiments confirmed lower mean MAEs on PEMS04 (p = 0.028) and PEMS08 (p = 0.042), although the corresponding RMSE differences did not reach the 0.05 significance threshold. Ablation, gate-weight, sensitivity, complexity and convergence analyses showed that temporal context was the main source of the improvement and that the gate provided a model-internal view of order selection with moderate overhead. CD-MRFG remains less accurate than several stronger recent baselines, so its value is a bounded and interpretable extension of ASTGCN rather than a universal state-of-the-art replacement.

1. Introduction

Short-term traffic flow forecasting supports route guidance, congestion warning, signal control and network-level traffic management. The task is difficult because detector observations form multivariate time series on a non-Euclidean road graph, and their evolution reflects local persistence, spatial propagation and recurring temporal regimes. Recurrent models represent temporal dependence [1,2], while graph-based models explicitly describe interactions among sensors [3,4,5,6]. A useful forecasting model must, therefore, balance temporal dynamics, graph structure and operating context without obscuring the source of its gains.
Graph-based traffic forecasting has progressed from fixed-topology convolution and diffusion [4,5] to adaptive graph learning and node-specific parameterization [6,7,8]. Attention and Transformer architectures further capture long-range or time-varying dependencies through global attention, adaptive embeddings and propagation-delay modelling [9,10,11,12]. These methods achieve strong predictive accuracy, but their adaptation usually changes the graph, node representation, attention map or latent embedding. The contribution of individual polynomial graph-convolution orders is seldom exposed as a context-conditioned quantity.
This distinction is relevant to ASTGCN. The model learns temporal and spatial attention from historical flow, but calendar semantics such as hour of day, day of week and weekend status remain implicit. Moreover, its Chebyshev graph convolution combines polynomial orders without an explicit sample-level mechanism that links their relative contributions to temporal context. The same receptive-field mixture can, therefore, be applied to peak, off-peak and weekend conditions even when their spatial dependence patterns differ.
CD-MRFG addresses this narrow limitation while retaining the ASTGCN backbone. Its first contribution is a five-dimensional calendar-context encoding that separates periodic semantics from raw flow measurements. Its second contribution is a block-specific Softmax gate that converts the encoded context into weights for different Chebyshev orders. Its third contribution is an evaluation on PEMS03, PEMS04 and PEMS08 that combines reproduced baselines, module ablation, three-seed tests, gate diagnostics, complexity profiling, hyperparameter sensitivity and quantitative convergence analysis. The resulting claim is deliberately bounded: the method improves its direct ASTGCN baseline and exposes context-dependent order allocation, but it does not outperform every recent model.
The remainder of this paper is organized as follows. Section 2 reviews temporal, graph-based and context-aware forecasting methods and distinguishes CD-MRFG from related adaptive models. Section 3 defines the forecasting task and notation. Section 4 presents the model architecture and training objective. Section 5 reports the comparative, robustness, ablation, interpretation, efficiency, sensitivity and convergence results. Section 6 summarizes the findings and their limitations.

2. Related Work

2.1. Temporal and Statistical Traffic Forecasting

Classical approaches such as historical averages and autoregressive models are transparent but rely on stationary or linear assumptions. LSTM [1] and GRU [2] relax these assumptions through gated state transitions and capture nonlinear temporal dependence. Their main limitation for network forecasting is structural: each sensor is primarily treated as part of a sequence, so road-network propagation must be inferred indirectly.

2.2. Graph-Based Spatio-Temporal Forecasting

Graph neural networks represent sensors as nodes connected by physical or learned relations. STGCN combines graph and temporal convolutions [4], DCRNN models directed diffusion with recurrent graph operators [5], and Graph WaveNet learns an adaptive adjacency matrix with dilated temporal convolutions [6]. ASTGCN adds spatial and temporal attention to Chebyshev graph convolution [3]. AGCRN learns node-adaptive parameters and a data-adaptive graph [7], STSGCN constructs localized synchronous spatio-temporal graphs [13], and MTGNN jointly learns a directed graph, mix-hop propagation and multi-scale temporal filters [8]. These models adapt topology, node parameters or propagation operators, but they do not directly condition each Chebyshev order on calendar context. Other graph-based extensions include optimized temporal–spatial gated graph convolution [14], T-GCN [15], and spatial–temporal position-aware graph convolution [16]. MRA-BGCN is particularly relevant because it uses multi-range attention to learn the importance of different neighbourhood ranges from node- and edge-interaction features [17]. Unlike MRA-BGCN, CD-MRFG retains the ASTGCN backbone and conditions Chebyshev-order weights explicitly on an external calendar-context vector.

2.3. Context-Aware and Adaptive Traffic Modelling

Context-aware and attention-based methods model variation beyond fixed local graph convolution. GMAN links historical and future representations through spatio-temporal and transform attention [10]. STID uses spatial and temporal identity embeddings without explicit message passing [18]. MSGFormer combines multi-scale temporal modelling with gated Transformer components [9], while PDFormer introduces dynamic short- and long-range attention masks and propagation-delay-aware transformation [11]. STAEformer shows that adaptive spatio-temporal embeddings can make a vanilla Transformer highly competitive [12]. These approaches motivate context-dependent modelling, but their internal adaptive objects differ from the order-wise receptive-field allocation studied here. Traffic Transformer [19], ST-ResNet [20], and STDN [21] further demonstrate how periodicity, residual spatio-temporal structure, and dynamically changing dependencies can be incorporated into traffic prediction.

2.4. Differences from Existing Methods

CD-MRFG does not replace adaptive graph learning or Transformer attention. It retains the predefined road graph and ASTGCN feature extractor, then conditions the contribution of self, first-order and higher-order Chebyshev components on a five-dimensional calendar vector. Table 1 makes this distinction explicit. The method is, therefore, closest to a controlled, interpretable modification of ASTGCN, whereas Graph WaveNet, AGCRN and MTGNN adapt graph relations, STID adapts identity embeddings, and recent Transformers adapt attention or latent spatio-temporal embeddings.

3. Problem Formulation

3.1. Traffic Network Graph

In the traffic forecasting task, the road network is represented as a weighted graph G = (V, E, A), where V is the set of sensor nodes, E denotes road connectivity and A is the adjacency matrix. PEMS03, PEMS04 and PEMS08 contain 358, 307 and 170 sensors, respectively, and all flow records are aggregated at 5 min intervals.
X t = x t 1 , , x t N T R N × C
Here, X t denotes the traffic state of all nodes at time step t, and C is the feature dimension. This study uses only traffic flow, so C = 1.

3.2. Multi-Step Traffic Forecasting Task

Given a sequence of L historical traffic-flow steps X t L + 1 : t , the graph G and the temporal context vector c t , the forecasting model learns a nonlinear mapping f_Theta to predict the next P time steps. Here, X t L + 1 : t belongs to R L × N × C , the predicted sequence belongs to R P × N × C , and both L and P are set to 12 in all experiments. Therefore, the model uses the previous 1 h to forecast the next 1 h.
X ^ t + 1 : t + P = f θ X t L + 1 : t , G , c t

3.3. Temporal Context Features

Traffic flow exhibits strong periodicity. The same sensor may follow different patterns across hours, weekdays and weekends. To represent this temporal context explicitly, a five-dimensional context vector is constructed at the forecasting origin:
c t = s h , q h , s d , q d , I weekend T R 5
s h = s i n 2 π h t 24 ,   q h = c o s 2 π h t 24
s d = s i n 2 π d t 7 ,   q d = c o s 2 π d t 7
where h t denotes the hour-of-day index in one day, d t denotes the day-of-week index in one week, and I w e e k e n d is the weekend indicator. The sine and cosine terms are used because hour and weekday are periodic variables.

3.4. Receptive Fields in Graph Convolution

Chebyshev graph convolution approximates a spectral graph filter with K polynomial orders. The k-th order can be interpreted as aggregating information from the k-hop neighborhood. The Chebyshev bases are defined recursively in Equations (6) and (7), where T k ( L ~ ) denotes the k-th Chebyshev polynomial and L ~ is the scaled graph Laplacian. Different orders, therefore, correspond to different spatial receptive fields, whose contributions need not be identical for short-term forecasting [22,23].
T 0 L ~ = I N ,   T 1 L ~ = L ~
T k L ~ = 2 L ~ T k 1 L ~ T k 2 L ~ ,   k 2

3.5. Evaluation Metrics

To evaluate multi-step traffic forecasting performance, this study uses mean absolute error (MAE), root mean square error (RMSE) and mean absolute percentage error (MAPE). MAE measures the average absolute deviation between predicted and observed traffic flows and provides a direct indication of overall forecasting accuracy. RMSE assigns larger penalties to large deviations and is, therefore, more sensitive to abrupt prediction errors around peak changes or congestion transitions. MAPE measures relative prediction error and is useful for comparing errors across different flow scales. Because traffic-flow records may contain zero or near-zero values, MAPE is calculated only on the valid set Ω, where the ground-truth value is larger than a small threshold. The three metrics are defined as follows:
M A E = 1 M i = 1 M y i y ^ i
R M S E = 1 M i = 1 M y i y ^ i 2
M A P E = 1 Ω i Ω y i y ^ i y i × 100 %
For all three metrics, lower values indicate better forecasting performance. MAPE is reported as a decimal fraction throughout the paper.

4. CD-MRFG: Context-Driven Dynamic Gating for ASTGCN

As shown in Figure 1, CD-MRFG consists of an ASTGCN spatio-temporal backbone, a context feature embedding module, a dynamic multi-order receptive-field gating module and a forecasting head. Historical traffic flow is processed by temporal attention, spatial attention, Chebyshev graph convolution and temporal convolution. The temporal-context branch explicitly displays the cyclic hour-of-day and day-of-week encodings rather than empirical traffic-flow profiles. The weekend indicator is represented as a binary variable, with zero for weekdays and one for weekend days. The resulting context representation generates sample-dependent Chebyshev-order weights for each ASTGCN block, allowing the model to adjust the spatial propagation range according to the temporal operating condition.

4.1. ASTGCN Spatio-Temporal Backbone

The ASTGCN backbone contains temporal attention, spatial attention, Chebyshev graph convolution, temporal convolution and residual connections [1]. Temporal attention selects informative historical steps, whereas spatial attention produces a block-specific matrix S b . Equation (11) gives the standard Chebyshev convolution. In CD-MRFG, the normalized spatial attention matrix is denoted by S b ^ , the order-specific propagation operator by M b , k , and the corresponding feature by Z b , k , as defined in Equations (14)–(16).
C h e b C o n v X = k = 0 K 1 T k L ~ X Θ k

4.2. Context Feature Embedding

The context embedding module maps c t to a sample-level hidden vector h c through a multilayer perceptron. This representation does not replace the traffic-flow input; instead, it provides conditional information for the gating network and captures differences in traffic propagation patterns under different temporal contexts.
h c = φ c t = M L P c t

4.3. Dynamic Multi-Order Receptive-Field Gating

For the b-th ASTGCN block, the gating network generates a K-dimensional weight vector g b from h c . After softmax normalization, the weights sum to one, allowing the model to select an appropriate spatial receptive field for each temporal context.
The motivation of the proposed gating mechanism is that different temporal contexts may require different spatial receptive fields. During peak hours, congestion may propagate to upstream and downstream sensors, so higher-order neighbourhoods may provide useful information. During off-peak periods, traffic states may be more locally stable, so lower-order neighbourhoods or self-information can be more reliable. Fixed Chebyshev-order weighting, therefore, restricts the modelling flexibility of graph convolution. By conditioning the order weights on temporal context, CD-MRFG provides a simple mechanism for representing context-dependent propagation ranges in a traffic network.
g b = S o f t m a x W g h c + b g , k = 0 K 1   g b , k = 1
M b , k = T k ( L ~ ) S b ^
Z b , k = M b , k T X b Θ b , k
Z b = k = 0 K 1 g b , k Z b , k
Figure 2 details this modelling mechanism. The five-dimensional temporal context vector is first encoded as a hidden representation. The gate then produces sample-dependent K-order weights for the b-th ASTGCN block. Each Chebyshev order corresponds to self-information, local neighbourhood propagation or more distant neighbourhood propagation. The weighted sum forms a gated graph-convolution output, which is combined with temporal convolution, residual connection and layer normalization to update spatio-temporal features.
Algorithm 1 summarizes the dynamic gated graph convolution in the b-th ASTGCN block. The procedure keeps the original spatial-attention graph-convolution framework but adds a context-conditioned modelling layer that controls how different Chebyshev basis functions contribute to the simulated traffic propagation process.
Algorithm 1. Context-conditioned dynamic multi-order receptive-field gating
Input: traffic-flow feature Xb, temporal context vector ct, Chebyshev polynomial set { T k ( L ~ ) } , and spatial attention matrix Sb
Output: gated graph-convolution feature Zb of the b-th ASTGCN block
1.hc ← MLPc(ct)
2.Ab ← MLPg,b(hc)
3.gb ← Softmax(Ab)
4. g b = [ g b , 0 , g b , 1 , , g b , K 1 ] , k = 0 K 1 g b , k = 1
5.Normalize Sb to obtain S b ^
6.Initialize Zb ← 0
7.for k = 0, 1, …, K − 1 do
8.  if k = 0 then
9.      Mb,kI
10.  else
11. M b , k T k ( L ~ ) S b ^
12.  end if
13.  Ub,k ← GraphConv(Xb, Mb,k, Θb,k)
14.  ZbZb + gb,k Ub,k
15.end for
16.Zb ← ReLU(Zb + βb)
17.return Zb

4.4. Training Objective

After the forecasting head produces the predicted sequence, CD-MRFG is optimized by minimizing the mean squared error between the predicted and observed traffic flows. Let y i and y ^ i denote the ground-truth and predicted values over all valid nodes, horizons and samples. The residual term and training loss are defined in Equations (17) and (18).
e i = y i y ^ i
L = 1 M i = 1 M e i 2

5. Experiments and Results

Experiments were conducted on PEMS03, PEMS04 and PEMS08. The evaluation addressed six questions: whether CD-MRFG improves its direct ASTGCN baseline, whether the result generalizes to a third sensor network, how it compares with recent strong baselines, whether the gains persist across random seeds, which components and hyperparameters control performance, and whether the added mechanism has acceptable computational and optimization costs. Unless otherwise stated, all reproduced models used the same chronological split, 12-step input, 12-step target and test metrics.

5.1. Datasets

PEMS03, PEMS04 and PEMS08 are public California freeway detector datasets sampled every 5 min. As shown in Table 2, PEMS03 contains 358 sensors and 26,208 time steps from 1 September to 30 November 2018. PEMS04 contains 307 sensors and 16,992 time steps from 1 January to 28 February 2018, and PEMS08 contains 170 sensors and 17,856 time steps from 1 July to 31 August 2016. Each dataset was divided chronologically into training, validation and test sets using a 6:2:2 ratio. Twelve historical observations were used to forecast the following twelve observations.

5.2. Experimental Settings and Metrics

Experiments were run on a workstation with a 13th Gen Intel Core i7-13700K CPU and an NVIDIA GeForce RTX 4090D GPU. Architecture-specific objectives and validation-selected settings followed the corresponding implementations, while the data split, input length, forecast horizon and evaluation code were fixed. As shown in Table 3, ASTGCN and CD-MRFG used Adam, mean squared error, a learning rate of 0.001 and a batch size of 32. Their common configuration contained two spatio-temporal blocks, Chebyshev order K = 3, 64 graph and temporal filters, and a CD-MRFG context hidden dimension of 64. The controlled benchmark runs of CD-MRFG achieved an overall MAE and RMSE of 16.98 and 28.59 on PEMS03, 20.82 and 32.77 on PEMS04, and 17.24 and 26.84 on PEMS08. Repeated-seed statistics are reported separately in Section 5.4.
Following common practice in traffic forecasting, mean absolute error (MAE), root mean square error (RMSE) and masked mean absolute percentage error (MAPE) are used as metrics. Their formal definitions are given in Equations (8)–(10). Let Y and Ŷ denote the ground-truth and predicted traffic flows, N the number of sensors and T the prediction horizon. Because traffic flow may contain zero or near-zero values, MAPE is calculated only on the valid set Ω, where the ground-truth value is larger than ε. In this paper, ε is set to 1 × 10−5. A lower MAE and RMSE indicate smaller forecasting error. RMSE is more sensitive to large errors, whereas MAPE measures relative error.

5.3. Comparison with Baseline Methods

The comparison includes HA, LSTM, GRU, DCRNN, STGCN, ASTGCN, AGCRN and STSGCN. All values in Table 4 were reproduced with the common chronological split and evaluation code. Model-specific architecture and optimization settings followed their implementations and were selected using validation data. ASTGCN is the primary controlled baseline because CD-MRFG modifies its graph-convolution stage. STID and MTGNN are reported separately in Table 5 as recent strong baselines on PEMS04 and PEMS08.
Errors increased with the prediction horizon for the learned baselines on all three datasets. On PEMS03, CD-MRFG achieved an overall MAE and RMSE of 16.98 and 28.59. It improved DCRNN by 5.67% in MAE and 1.68% in RMSE, STGCN by 15.48% and 12.00%, and STSGCN by 4.34% and 2.46%, respectively. AGCRN remained more accurate, with an overall MAE and RMSE of 15.51 and 27.17. The third dataset, therefore, supports generalization beyond PEMS04 and PEMS08 while preserving the bounded accuracy claim.
On PEMS04, CD-MRFG reduced overall MAE and RMSE by 6.13% and 4.25% relative to DCRNN, and by 20.93% and 19.32% relative to STGCN. On PEMS08, it reduced MAE by 0.52% relative to DCRNN, whereas its RMSE was 0.60% higher. AGCRN produced lower overall errors than CD-MRFG on PEMS04 and PEMS08. These mixed results show that the proposed gate improves an ASTGCN-style model but does not dominate adaptive graph learning.
The direct ASTGCN comparison was consistent across datasets. CD-MRFG reduced overall MAE and RMSE from 18.66 and 31.05 to 16.98 and 28.59 on PEMS03, from 22.79 and 35.02 to 20.82 and 32.77 on PEMS04, and from 18.88 and 28.83 to 17.24 and 26.84 on PEMS08. The corresponding MAE reductions were 9.00%, 8.64% and 8.69%. This controlled comparison supports the specific claim that explicit calendar context and order gating improve the ASTGCN backbone.
Table 6 provides a stricter comparison with STID and MTGNN. Both recent baselines achieved lower overall errors than CD-MRFG on PEMS04 and PEMS08. The proposed method is, therefore, evaluated for its controlled gain over ASTGCN, explicit order-level diagnostics and modest architectural change, not for state-of-the-art accuracy.
MTGNN and STID benefit from learned graph structure or strong identity embeddings and, therefore, define an important accuracy boundary. Their advantage does not invalidate the controlled ASTGCN comparison, but it prevents a universal superiority claim.

5.4. Robustness Across Random Seeds

ASTGCN and CD-MRFG were independently trained with seeds 1, 42 and 2026 on PEMS04 and PEMS08. Table 5 reports mean ± standard deviation over the three runs. Paired two-sided t-tests compare the two models using matched seeds. Given the small sample size, the p-values are interpreted together with the paired differences rather than as standalone proof.
CD-MRFG reduced the mean MAE for every matched seed. The MAE difference was significant on PEMS04 (p = 0.028) and PEMS08 (p = 0.042). The mean RMSE was also lower, but the paired tests did not cross the 0.05 threshold on PEMS04 (p = 0.181) or PEMS08 (p = 0.057). MAPE differences were not significant. These results support a reproducible MAE gain while placing an explicit statistical boundary on the RMSE and MAPE claims.

5.5. Ablation Study

Four variants are evaluated on PEMS04 and PEMS08 to isolate the effect of each component. M1 is the original ASTGCN baseline. M2 adds fixed multi-order graph-convolution gating to ASTGCN. M3 uses only the context feature embedding module. M4, namely CD-MRFG, combines context embedding with dynamic multi-order receptive-field gating. Table 7 reports the average 12-step results.
  • M1 ASTGCN. This baseline measures the forecasting ability of the original spatio-temporal attention graph convolution without temporal context or multi-order gating. Its average MAE is 22.79 and RMSE is 35.02 on PEMS04, while its average MAE is 18.88 and RMSE is 28.83 on PEMS08. These results show that ASTGCN captures basic spatio-temporal dependencies, but its input remains dominated by historical flow values and does not explicitly encode temporal context.
  • M2 fixed multi-order graph-convolution gating. M2 does not provide a neutral modification to the ASTGCN baseline. On PEMS04, its MAE remains nearly unchanged, decreasing from 22.79 to 22.78, whereas its RMSE increases from 35.02 to 36.45, corresponding to a relative deterioration of 4.08%. Because RMSE assigns greater weight to large residuals, this divergence indicates that fixed order weighting increases the magnitude of difficult forecasting errors even when the average absolute error remains stable. A fixed gate applies the same receptive-field mixture to all samples and cannot distinguish traffic states dominated by local variation from states involving broader spatial propagation. This mismatch can cause excessive neighbourhood aggregation in local regimes or insufficient higher-order propagation during spatially extended congestion. On PEMS08, M2 produces only marginal reductions in MAE and RMSE, from 18.88 and 28.83 to 18.85 and 28.56, respectively, while MAPE increases from 0.12 to 0.13. The opposite RMSE changes across datasets indicate that a fixed receptive-field mixture interacts with dataset-specific topology and traffic-regime composition rather than providing a robust improvement. In contrast, M4 conditions the order weights on temporal context and reduces the PEMS04 RMSE from 36.45 to 32.77 and the PEMS08 RMSE from 28.56 to 26.84 relative to M2. These results indicate that the effectiveness of multi-order gating depends on sample-specific contextual modulation rather than on the introduction of additional order weights alone.
  • M3 context feature embedding. This variant evaluates the contribution of temporal context itself. M3 reduces MAE to 21.34 and RMSE to 33.81 on PEMS04 and reduces MAE to 17.83 and RMSE to 27.25 on PEMS08, clearly outperforming M1 on both datasets. Hour, weekday and weekend indicators, therefore, provide useful periodic semantics that help distinguish morning and evening peaks, off-peak periods and weekend traffic.
  • M4 CD-MRFG. The full model combines context embedding with dynamic multi-order receptive-field gating and achieves the best results among the four variants. It obtains an MAE of 20.82, RMSE of 32.77 and MAPE of 0.14 on PEMS04, and an MAE of 17.24 and RMSE of 26.84 on PEMS08. Its further improvement over M3 indicates that context is useful not only as an auxiliary input representation but also as a condition for selecting graph-convolution orders.

5.6. Gate Weight Analysis Under Different Traffic Periods

Gate weights were extracted from the PEMS04 and PEMS08 test sets to determine whether order allocation changed with temporal context. Table 8 reports the mean weights assigned to orders 0, 1 and 2 in the two ASTGCN blocks for four traffic periods. Each row sums to one because the gate uses a Softmax output.
The distributions varied across traffic periods and blocks. On PEMS04, block 1 emphasized order 2 during the morning peak and daytime off-peak periods, whereas block 2 emphasized order 1 in the morning peak and order 0 in the evening peak and at night. On PEMS08, block 1 mainly balanced orders 0 and 1, while block 2 assigned larger weights to order 2 in the three daytime periods. The gate, therefore, did not collapse to a single fixed mixture.
These weights are model-internal allocation coefficients, not measurements of physical traffic propagation or causal influence. Their variation supports the narrower conclusion that CD-MRFG uses different polynomial graph components under different calendar contexts. Physical interpretation would require external traffic-state evidence or controlled interventions.

5.7. Model Interpretation and Applicability Boundary

CD-MRFG is designed for short-term forecasting on fixed sensor graphs with regular daily and weekly regimes. The gate converts calendar context into a convex combination of self, first-order and higher-order graph features. This construction makes the model’s receptive-field preference inspectable without replacing the ASTGCN backbone.
The applicability boundary is equally important. The five context variables do not represent holidays, incidents, weather, roadworks or special events. The method may, therefore, be less reliable when non-recurring disturbances dominate the traffic state. It also retains a predefined graph and cannot recover missing or evolving connectivity as directly as adaptive graph-learning models.
AGCRN, STID and MTGNN achieved lower absolute errors in several comparisons. CD-MRFG should, thus, be used when the objective is a controlled ASTGCN improvement with explicit order-level diagnostics and modest structural change, rather than when forecasting accuracy alone determines model choice.

5.8. Complexity and Computational Overhead

Because CD-MRFG adds only a context encoder and a dynamic receptive-field gate to ASTGCN, its overhead mainly comes from lightweight multilayer perceptrons and scalar order-wise weighting. Table 9 reports profiling under a batch size of 32 on CUDA. On PEMS04, the parameter count increases from 0.450 M to 0.459 M, and latency increases from 24.06 ms to 27.86 ms. On PEMS08, the parameter count increases from 0.179 M to 0.189 M, and latency increases from 17.94 ms to 19.53 ms. The measured PEMS08 checkpoint sizes are 0.6966 MB for ASTGCN and 0.7342 MB for CD-MRFG. The measured values indicate limited model-size growth and moderate inference overhead.

5.9. Hyperparameter Sensitivity

A one-factor-at-a-time analysis was conducted on PEMS04 with seed 42. Chebyshev order K, context hidden dimension and the number of spatio-temporal blocks were varied while the remaining settings were fixed. The analysis was performed after the main benchmark and was not used to retrospectively select the reported test configuration. Table 10 reports PEMS04 hyperparameter sensitivity results and Figure 3 demonstrates PEMS04 sensitivity to the varied parameters.
Increasing K from 2 to 4 progressively reduced MAE and RMSE, but K = 4 increased training time by 11.1% relative to K = 3. A hidden dimension of 32 was more accurate and smaller than the default 64, whereas 128 degraded both errors, indicating that a larger context encoder is not automatically beneficial. Three blocks achieved the lowest errors but increased parameters and training time by about 52% and 50% relative to two blocks. The main setting, therefore, represents a fixed benchmark configuration rather than the best point from post hoc tuning.

5.10. Quantitative Convergence Analysis

Table 11 reports quantitative convergence results from the completed three-seed logs. E90 and E95 denote the first epochs that reached 90% and 95% of the final achieved validation-loss reduction. In Figure 4, normalized area under the validation curve (AUC) summarizes the full optimization trajectory, with a smaller value indicating faster loss reduction. Late-stage variability is measured by the coefficient of variation over the final ten epochs.
Both models reached E90 within approximately five epochs. CD-MRFG had a lower normalized AUC on both datasets. The paired difference was significant on PEMS08 (p = 0.031) and marginal on PEMS04 (p = 0.054). On PEMS08, CD-MRFG also reduced late-stage variability from 3.18% to 0.84%. These optimization benefits required additional time: the mean training duration increased by 4.58 min on PEMS04 and 4.16 min on PEMS08. The results support a more efficient loss trajectory, particularly on PEMS08, rather than a universal claim of faster threshold attainment.

5.11. Forecasting Visualization

In Figure 5 and Figure 6, representative PEMS04 and PEMS08 test sequences are plotted with ground-truth flow, ASTGCN predictions and CD-MRFG predictions. The curves provide qualitative evidence about turning-point response and trend continuity. PEMS03 is evaluated quantitatively in Table 4 but is not used to select an additional visualization.
On PEMS04, the selected sequence contains evident morning-like growth, a high-flow plateau, subsequent decline and recovery. These stages are useful for evaluating multi-step forecasting because errors easily accumulate around turning points. The ASTGCN baseline captures the broad temporal trend, but its predicted curve tends to lag when the flow changes quickly. CD-MRFG gives a smoother yet more responsive trajectory, especially around the falling segment after the local peak and the later recovery interval.
The PEMS08 visualization provides a second case with a different sensor scale and temporal distribution. Although the overall profile differs from PEMS04, the same phenomenon can be observed: baseline predictions are delayed in several peak-decline and trough-recovery intervals, whereas CD-MRFG remains closer to the ground-truth direction of change. This cross-dataset consistency supports the argument that the proposed context-conditioned gate improves the dynamic selection of spatial receptive fields instead of overfitting one particular dataset.
The visual results complement the aggregate errors in Table 4. They show that CD-MRFG follows several rise, decline and recovery segments more closely than ASTGCN. These examples are illustrative rather than statistical evidence and should be interpreted together with the full test-set metrics and repeated-seed analysis.

6. Conclusions

This paper presented CD-MRFG, a context-gated ASTGCN extension that maps calendar information to Chebyshev-order weights. Controlled experiments on PEMS03, PEMS04 and PEMS08 consistently improved the direct ASTGCN baseline, and three-seed tests confirmed significant mean MAE reductions on PEMS04 and PEMS08. Ablation identified temporal context as the main contributor, while gate diagnostics showed non-constant order allocation across periods and blocks. Sensitivity and convergence analyses further exposed the effects of model scale and the associated optimization cost. CD-MRFG did not outperform AGCRN, STID or MTGNN in every setting, and its gate weights are not causal measurements of traffic propagation. The method is, therefore, most appropriate as a lightweight and inspectable improvement to ASTGCN on fixed sensor graphs with recurring calendar regimes.

Author Contributions

Methodology, Y.Z.; software, J.L.; validation, Z.Y., Z.D. and Y.K.; data curation, Z.D.; writing—original draft preparation, J.L.; writing—review and editing, Y.Z.; visualization, Y.Z.; supervision, Y.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by National Natural Science Foundation of China, grant number 52302408.

Data Availability Statement

The PEMS03, PEMS04 and PEMS08 datasets used in this study are public traffic forecasting benchmarks. The processed splits follow the chronological 6:2:2 protocol described in Section 5.1.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ASTGCNAttention-based Spatial–Temporal Graph Convolutional Network
CD-MRFGContext-Driven Dynamic Multi-Order Receptive-Field Gated ASTGCN
PEMS04Performance Measurement System Dataset 04
PEMS08Performance Measurement System Dataset 08
HAHistorical Average
LSTMLong Short-Term Memory
GRUGated Recurrent Unit
GCNGraph Convolutional Network
STGCNSpatio-Temporal Graph Convolutional Network
DCRNNDiffusion Convolutional Recurrent Neural Network
AGCRNAdaptive Graph Convolutional Recurrent Network
STSGCNSpatial–Temporal Synchronous Graph Convolutional Network
STIDSpatial–Temporal Identity
GMANGraph Multi-Attention Network
STDNSpatial–Temporal Dynamic Network
T-GCNTemporal Graph Convolutional Network
PEMS03Performance Measurement System Dataset 03
MTGNNMultivariate Time-Series Graph Neural Network
PDFormerPropagation Delay-Aware Dynamic Long-Range Transformer
STAEformerSpatio-Temporal Adaptive Embedding Transformer
AUCArea Under the Curve
MRA-BGCNMulti-Range Attentive Bicomponent Graph Convolutional Network

References

  1. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv 2014, arXiv:1412.3555. [Google Scholar] [CrossRef] [Scilit]
  3. Guo, S.; Lin, Y.; Feng, N.; Song, C.; Wan, H. Attention Based Spatial-Temporal Graph Convolutional Networks for Traffic Flow Forecasting. Proc. AAAI Conf. Artif. Intell. 2019, 33, 922–929. [Google Scholar] [CrossRef] [Scilit]
  4. Yu, B.; Yin, H.; Zhu, Z. Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden, 13–19 July 2018; AAAI Press: Washington, DC, USA, 2018; pp. 3634–3640. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Li, Y.; Yu, R.; Shahabi, C.; Liu, Y. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018; ICLR: Appleton, WI, USA, 2018; pp. 1–16. [Google Scholar] [CrossRef] [Scilit]
  6. Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Zhang, C. Graph WaveNet for Deep Spatio-Temporal Graph Modeling. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, Macau, China, 10–16 August 2019; AAAI Press: Washington, DC, USA, 2019; pp. 1907–1913. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Bai, L.; Yao, L.; Li, C.; Wang, X.; Wang, C. Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting. Adv. Neural Inf. Process. Syst. 2020, 33, 17804–17815. [Google Scholar] [CrossRef] [Scilit]
  8. Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Chang, X.; Zhang, C. Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Virtual Event, CA, USA, 6–10 July 2020; ACM: New York, NY, USA, 2020; pp. 753–763. [Google Scholar] [CrossRef] [Scilit]
  9. Li, W.; Chen, J.; Zhang, Y.; Sun, R.; Xia, S.; Pan, Z.; Luo, J. MSGFormer: Revolutionizing Traffic Flow Prediction with Multiscale and Gated Transformer Architecture. IEEE Internet Things J. 2025, 12, 2014–2025. [Google Scholar] [CrossRef] [Scilit]
  10. Zheng, C.P.; Fan, X.L.; Wang, C.; Qi, J.Z. GMAN: A Graph Multi-Attention Network for Traffic Prediction. Proc. AAAI Conf. Artif. Intell. 2020, 34, 1234–1241. [Google Scholar] [CrossRef] [Scilit]
  11. Jiang, J.; Han, C.; Zhao, W.X.; Wang, J. PDFormer: Propagation Delay-Aware Dynamic Long-Range Transformer for Traffic Flow Prediction. Proc. AAAI Conf. Artif. Intell. 2023, 37, 4365–4373. [Google Scholar] [CrossRef] [Scilit]
  12. Liu, H.; Dong, Z.; Jiang, R.; Deng, J.; Deng, J.; Chen, Q.; Song, X. Spatio-Temporal Adaptive Embedding Makes Vanilla Transformer SOTA for Traffic Forecasting. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, Birmingham, UK, 21–25 October 2023; ACM: New York, NY, USA, 2023; pp. 4125–4129. [Google Scholar] [CrossRef] [Scilit]
  13. Song, C.; Lin, Y.; Guo, S.; Wan, H. Spatial-Temporal Synchronous Graph Convolutional Networks: A New Framework for Spatial-Temporal Network Data Forecasting. Proc. AAAI Conf. Artif. Intell. 2020, 34, 914–921. [Google Scholar] [CrossRef] [Scilit]
  14. Guo, K.; Hu, Y.; Sun, Y.; Qian, Z.S.; Gao, J.; Yin, B. An Optimized Temporal-Spatial Gated Graph Convolution Network for Traffic Forecasting. IEEE Intell. Transp. Syst. Mag. 2022, 14, 153–162. [Google Scholar] [CrossRef]
  15. Zhao, L.; Song, Y.; Zhang, C.; Liu, Y.; Wang, P.; Lin, T.; Deng, M.; Li, H. T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction. IEEE Trans. Intell. Transp. 2020, 21, 3848–3858. [Google Scholar] [CrossRef] [Scilit]
  16. Zhao, Y.; Lin, Y.; Wen, H.; Wei, T.; Jin, X.; Wan, H. Spatial-Temporal Position-Aware Graph Convolution Networks for Traffic Flow Forecasting. IEEE Trans. Intell. Transp. Syst. 2023, 24, 8650–8666. [Google Scholar] [CrossRef] [Scilit]
  17. Chen, W.; Chen, L.; Xie, Y.; Cao, W.; Gao, Y.; Feng, X. Multi-Range Attentive Bicomponent Graph Convolutional Network for Traffic Forecasting. Proc. AAAI Conf. Artif. Intell. 2020, 34, 3529–3536. [Google Scholar] [CrossRef] [Scilit]
  18. Shao, Z.; Zhang, Z.; Wang, F.; Wei, W.; Xu, Y. Spatial-Temporal Identity: A Simple yet Effective Baseline for Multivariate Time Series Forecasting. arXiv 2022, arXiv:2208.05233. [Google Scholar] [CrossRef] [Scilit]
  19. Cai, L.; Janowicz, K.; Mai, G.; Yan, B.; Zhu, R. Traffic Transformer: Capturing the Continuity and Periodicity of Time Series for Traffic Forecasting. Trans. GIS 2020, 24, 736–755. [Google Scholar] [CrossRef] [Scilit]
  20. Zhang, J.B.; Zheng, Y.; Qi, D.K. Deep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction. Proc. AAAI Conf. Artif. Intell. 2017, 31, 1655–1661. [Google Scholar] [CrossRef] [Scilit]
  21. Yao, H.X.; Tang, X.F.; Wei, H.; Zheng, G.; Li, Z. Revisiting Spatial-Temporal Similarity: A Deep Learning Framework for Traffic Prediction. Proc. AAAI Conf. Artif. Intell. 2019, 33, 5668–5675. [Google Scholar] [CrossRef] [Scilit]
  22. Defferrard, M.; Bresson, X.; Vandergheynst, P. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In Advances in Neural Information Processing Systems; NeurIPS: Barcelona, Spain, 2016; pp. 3844–3852. [Google Scholar] [CrossRef] [Scilit]
  23. Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the International Conference on Learning Representations; ICLR: Toulon, France, 2017; pp. 1–14. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overall architecture of CD-MRFG.
Figure 1. Overall architecture of CD-MRFG.
Modelling 07 00157 g001
Figure 2. Context-driven dynamic multi-order receptive-field gating.
Figure 2. Context-driven dynamic multi-order receptive-field gating.
Modelling 07 00157 g002
Figure 3. PEMS04 sensitivity to Chebyshev order, context hidden dimension and number of spatio-temporal blocks. Orange markers denote the fixed benchmark settings.
Figure 3. PEMS04 sensitivity to Chebyshev order, context hidden dimension and number of spatio-temporal blocks. Orange markers denote the fixed benchmark settings.
Modelling 07 00157 g003
Figure 4. Normalized validation-loss trajectories across three seeds on (a) PEMS04 and (b) PEMS08. Shaded regions denote mean ± one standard deviation.
Figure 4. Normalized validation-loss trajectories across three seeds on (a) PEMS04 and (b) PEMS08. Shaded regions denote mean ± one standard deviation.
Modelling 07 00157 g004
Figure 5. Forecasting visualization on PEMS04.
Figure 5. Forecasting visualization on PEMS04.
Modelling 07 00157 g005
Figure 6. Forecasting visualization on PEMS08.
Figure 6. Forecasting visualization on PEMS08.
Modelling 07 00157 g006
Table 1. Methodological differences between CD-MRFG and representative adaptive traffic forecasting models.
Table 1. Methodological differences between CD-MRFG and representative adaptive traffic forecasting models.
ModelPrimary Adaptive ObjectGraph RepresentationContext or Scale MechanismDifference from CD-MRFG
ASTGCN [3]Spatial and temporal attentionPredefined road graphStatic combination of Chebyshev ordersDirect backbone; no explicit calendar-conditioned order gate
Graph WaveNet [6]Adaptive adjacencyLearned and predefined graphsDilated temporal convolutionLearns graph relations rather than calendar-conditioned order weights
AGCRN [7]Node embeddings and node-specific parametersData-adaptive graphAdaptive recurrent graph convolutionAdapts nodes and topology rather than explicit Chebyshev orders
GMAN [10]Spatio-temporal attentionAttention-derived dependenciesTransform attention between historical and future statesUses global attention instead of order-wise graph convolution
STID [18]Spatial and temporal identitiesNo explicit graph propagationIdentity embeddingsStrong embedding baseline without an interpretable graph-order gate
MTGNN [8]Directed graph and mix-hop propagationLearned graphDilated inception temporal moduleLearns topology and mix-hop features without calendar conditioning
Transformer models [9,11,12]Attention, adaptive embeddings or delay masksGlobal or dynamically masked relationsLong-range attention and multi-scale temporal modellingBroader architectures; no explicit ASTGCN Chebyshev-order allocation
CD-MRFGChebyshev-order contributionPredefined road graphFive-dimensional calendar context and Softmax order gateLightweight ASTGCN extension with inspectable order weights
Table 2. Dataset statistics.
Table 2. Dataset statistics.
DatasetNodesTime StepsTime Range
PEMS0335826,2081 September 2018 to 30 November 2018
PEMS0430716,9921 January 2018 to 28 February 2018
PEMS0817017,8561 July 2016 to 31 August 2016
Table 3. Experimental settings.
Table 3. Experimental settings.
DatasetLearning RateBatch SizeOptimizerInput StepsForecast StepsGraph OrderST BlocksContext Hidden Dimension
PEMS030.00132Adam12123264
PEMS040.00132Adam12123264
PEMS080.00132Adam12123264
Table 4. Forecasting performance of CD-MRFG and reproduced baselines on PEMS03, PEMS04 and PEMS08 datasets.
Table 4. Forecasting performance of CD-MRFG and reproduced baselines on PEMS03, PEMS04 and PEMS08 datasets.
DatasetModel5 min15 min30 min45 min60 minAverage
MAERMSEMAERMSEMAERMSEMAERMSEMAERMSEMAERMSE
PEMS03HA26.1747.5126.1747.5026.1647.5026.1647.5026.1547.4926.1647.50
LSTM13.3922.9816.1127.3919.3231.9522.6536.6326.2341.5019.8833.12
GRU13.4123.1216.1227.4819.3532.2822.6736.8926.3041.8519.9133.37
DCRNN12.7420.3115.2624.5917.7028.4520.1232.0122.5535.5018.0029.08
STGCN16.6127.0518.1129.4019.7631.9121.5334.4923.5837.4820.0932.49
ASTGCN13.4722.7815.8526.6018.2730.3020.7633.8823.5537.7018.6631.05
AGCRN13.1523.0914.4025.3315.6127.2916.3728.6117.1729.6015.5127.17
STSGCN14.2023.1815.9126.1217.5428.8319.1231.4620.8934.2317.7529.31
CD-MRFG13.2922.3215.2725.6516.8828.3518.3530.6619.9933.1216.9828.59
PEMS04HA26.0441.6626.0441.6626.0541.6726.0641.6726.0641.6726.0541.67
LSTM18.3328.9721.2233.1224.8938.2928.8643.6233.1749.4025.6639.71
GRU18.1928.8721.1533.0824.9138.3328.8843.7333.2249.5125.6539.75
DCRNN17.4027.5919.8830.9622.0633.9424.0736.5725.9339.0722.1834.23
STGCN22.4935.5923.8837.3625.7139.7028.0342.6430.8946.2826.3340.62
ASTGCN18.3028.7820.3431.6222.3834.4024.5737.2127.2340.6222.7935.02
AGCRN17.8628.6318.5929.9819.5631.6620.0732.4620.6333.3319.4431.44
STSGCN17.9628.4919.5430.9321.2133.5522.8435.8624.4538.1021.3833.81
CD-MRFG17.5227.9419.1930.3920.6532.5122.0134.3923.7736.7020.8232.77
PEMS08HA23.2740.5623.2640.5523.2540.5223.2340.5023.2240.4723.2440.52
LSTM13.9221.3816.4925.6419.5430.5622.6835.0026.3239.9620.1031.56
GRU13.9421.4316.5425.7419.6830.7522.9035.3426.5040.1520.2331.75
DCRNN13.3520.3515.6523.9617.2326.5218.8528.7520.2030.6817.3326.68
STGCN20.0630.3220.9831.6622.2633.5923.8135.7125.8538.4722.6834.18
ASTGCN15.3523.2817.1226.1919.1129.1821.0331.9123.4735.2218.8828.83
AGCRN13.7721.2414.7323.1515.7425.1016.6126.6817.3527.8115.7725.17
STSGCN14.1521.5815.7224.3617.1326.8818.5229.1719.9631.3117.2827.18
CD-MRFG13.9421.4315.7524.4817.1526.7418.3628.4620.0430.7217.2426.84
Table 5. Three-seed robustness and paired significance tests for ASTGCN and CD-MRFG.
Table 5. Three-seed robustness and paired significance tests for ASTGCN and CD-MRFG.
DatasetModelMAERMSEMAPE
PEMS04ASTGCN22.2984 ± 0.482934.5454 ± 0.64680.1563 ± 0.0054
PEMS04CD-MRFG21.7245 ± 0.383933.9865 ± 0.50870.1563 ± 0.0104
PEMS08ASTGCN18.8479 ± 0.768428.7259 ± 1.06940.1232 ± 0.0089
PEMS08CD-MRFG16.8782 ± 0.056126.2775 ± 0.12860.1151 ± 0.0020
Table 6. Overall comparison with reproduced recent strong baselines on PEMS04 and PEMS08.
Table 6. Overall comparison with reproduced recent strong baselines on PEMS04 and PEMS08.
DatasetModelRoleBest EpochMAERMSEMAPE
PEMS04CD-MRFGContext-gated ASTGCN extension2320.8232.770.14
STID [18]Identity-embedding baseline8618.4029.920.13
MTGNN [8]Adaptive graph and mix-hop baseline10019.0531.620.13
PEMS08CD-MRFGContext-gated ASTGCN extension5917.2426.840.11
STID [18]Identity-embedding baseline8914.2223.400.09
MTGNN [8]Adaptive graph and mix-hop baseline9815.4724.520.10
Table 7. Average 12-step results of the ablation variants on PEMS04 and PEMS08.
Table 7. Average 12-step results of the ablation variants on PEMS04 and PEMS08.
ModelPEMS04PEMS08
MAERMSEMAPEMAERMSEMAPE
M1 ASTGCN22.7935.020.1518.8828.830.12
M2 fixed multi-order gating22.7836.450.1518.8528.560.13
M3 context embedding21.3433.810.1417.8327.250.11
M4 CD-MRFG20.8232.770.1417.2426.840.11
Table 8. Mean dynamic gate weights under different traffic periods.
Table 8. Mean dynamic gate weights under different traffic periods.
DatasetTraffic PeriodSamplesBlockOrder 0Order 1Order 2
PEMS04Morning peak43210.4400.0360.524
PEMS04Morning peak43220.0100.7480.242
PEMS04Daytime off-peak100810.3520.0540.594
PEMS04Daytime off-peak100820.1070.3620.531
PEMS04Evening peak43210.4310.3550.214
PEMS04Evening peak43220.5760.0340.390
PEMS04Night152210.2850.3060.409
PEMS04Night152220.5190.2820.199
PEMS08Morning peak43210.5230.4760.002
PEMS08Morning peak43220.2480.0440.707
PEMS08Daytime off-peak104610.5060.4830.012
PEMS08Daytime off-peak104620.2910.0120.697
PEMS08Evening peak46810.4100.4680.123
PEMS08Evening peak46820.1450.2090.646
PEMS08Night162110.3610.5890.050
PEMS08Night162120.1920.3580.450
Note: Order 0 denotes self-information, order 1 denotes first-order neighbor propagation, and order 2 denotes second-order neighborhood propagation. Values are averaged over test samples in each traffic period.
Table 9. Parameter scale and computational profiling results.
Table 9. Parameter scale and computational profiling results.
DatasetModelParameters (M)Checkpoint Size (MB)Latency (ms per Batch)Throughput (Samples per Second)
PEMS04ASTGCN0.4500311.728724.05941330.04
PEMS04CD-MRFG0.4591251.766427.86371148.4482
PEMS08ASTGCN0.1794560.696617.94001783.7233
PEMS08CD-MRFG0.1885500.734219.53441638.1363
Table 10. PEMS04 hyperparameter sensitivity results.
Table 10. PEMS04 hyperparameter sensitivity results.
ParameterValueMAERMSEMAPEParametersTraining Time (min)
Chebyshev order K224.091136.90840.1863454 83529.90
Chebyshev order K3 (default)22.157534.49520.1681459 12534.70
Chebyshev order K421.571533.73880.1522463 41538.56
Context hidden dimension3221.483333.94540.1506452 53334.39
Context hidden dimension64 (default)22.157534.49520.1681459 12534.70
Context hidden dimension12823.566136.68470.1695484 59733.17
ST blocks122.972235.36280.1745220 35317.83
ST blocks2 (default)22.157534.49520.1681459 12534.70
ST blocks321.422133.88590.1405697 89752.15
Table 11. Quantitative convergence results over three random seeds.
Table 11. Quantitative convergence results over three random seeds.
DatasetModelE90E95Normalized AUCTail CV (%)Optimization Time (min)
PEMS04ASTGCN4.33 ± 2.317.67 ± 3.060.0222 ± 0.00391.25 ± 0.4231.13 ± 0.54
PEMS04CD-MRFG4.00 ± 1.008.00 ± 3.460.0190 ± 0.00251.21 ± 0.7235.71 ± 2.55
PEMS08ASTGCN4.67 ± 2.087.33 ± 2.520.0262 ± 0.00263.18 ± 1.3722.11 ± 0.10
PEMS08CD-MRFG3.33 ± 0.587.00 ± 2.000.0163 ± 0.00110.84 ± 0.3626.28 ± 0.47
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, Y.; Liang, J.; Yuan, Z.; Dai, Z.; Kang, Y. Context-Gated Graph Modelling for Traffic Flow Forecasting. Modelling 2026, 7, 157. https://doi.org/10.3390/modelling7040157

AMA Style

Zhang Y, Liang J, Yuan Z, Dai Z, Kang Y. Context-Gated Graph Modelling for Traffic Flow Forecasting. Modelling. 2026; 7(4):157. https://doi.org/10.3390/modelling7040157

Chicago/Turabian Style

Zhang, Yuzhuo, Jialin Liang, Ziqiong Yuan, Zanzan Dai, and Yaozheng Kang. 2026. "Context-Gated Graph Modelling for Traffic Flow Forecasting" Modelling 7, no. 4: 157. https://doi.org/10.3390/modelling7040157

APA Style

Zhang, Y., Liang, J., Yuan, Z., Dai, Z., & Kang, Y. (2026). Context-Gated Graph Modelling for Traffic Flow Forecasting. Modelling, 7(4), 157. https://doi.org/10.3390/modelling7040157

Article Metrics

Back to TopTop