Skip to Content
ModellingModelling
  • Article
  • Open Access

5 August 2026

18 Pages

Context-Gated Graph Modelling for Traffic Flow Forecasting

,
,
,
and
1
College of Vehicle and Traffic Engineering, Henan University of Science and Technology, Luoyang 471003, China
2
School of Transportation and Logistics Engineering, Wuhan University of Technology, Wuhan 430070, China
*
Author to whom correspondence should be addressed.

Abstract

Traffic states evolve on irregular sensor graphs and vary with calendar context, yet the original ASTGCN does not explicitly model how the contribution of different graph receptive fields changes across traffic periods. This paper proposes CD-MRFG, a context-gated extension of ASTGCN that encodes hour-of-day, day-of-week and weekend information and uses the resulting representation to weight Chebyshev graph-convolution orders in each spatio-temporal block. Under a common 12-step forecasting protocol, CD-MRFG reduced the overall MAE and RMSE of the reproduced ASTGCN baseline from 18.66 and 31.05 to 16.98 and 28.59 on PEMS03, from 22.79 and 35.02 to 20.82 and 32.77 on PEMS04, and from 18.88 and 28.83 to 17.24 and 26.84 on PEMS08. Three-seed experiments confirmed lower mean MAEs on PEMS04 (p = 0.028) and PEMS08 (p = 0.042), although the corresponding RMSE differences did not reach the 0.05 significance threshold. Ablation, gate-weight, sensitivity, complexity and convergence analyses showed that temporal context was the main source of the improvement and that the gate provided a model-internal view of order selection with moderate overhead. CD-MRFG remains less accurate than several stronger recent baselines, so its value is a bounded and interpretable extension of ASTGCN rather than a universal state-of-the-art replacement.

1. Introduction

Short-term traffic flow forecasting supports route guidance, congestion warning, signal control and network-level traffic management. The task is difficult because detector observations form multivariate time series on a non-Euclidean road graph, and their evolution reflects local persistence, spatial propagation and recurring temporal regimes. Recurrent models represent temporal dependence [1,2], while graph-based models explicitly describe interactions among sensors [3,4,5,6]. A useful forecasting model must, therefore, balance temporal dynamics, graph structure and operating context without obscuring the source of its gains.
Graph-based traffic forecasting has progressed from fixed-topology convolution and diffusion [4,5] to adaptive graph learning and node-specific parameterization [6,7,8]. Attention and Transformer architectures further capture long-range or time-varying dependencies through global attention, adaptive embeddings and propagation-delay modelling [9,10,11,12]. These methods achieve strong predictive accuracy, but their adaptation usually changes the graph, node representation, attention map or latent embedding. The contribution of individual polynomial graph-convolution orders is seldom exposed as a context-conditioned quantity.
This distinction is relevant to ASTGCN. The model learns temporal and spatial attention from historical flow, but calendar semantics such as hour of day, day of week and weekend status remain implicit. Moreover, its Chebyshev graph convolution combines polynomial orders without an explicit sample-level mechanism that links their relative contributions to temporal context. The same receptive-field mixture can, therefore, be applied to peak, off-peak and weekend conditions even when their spatial dependence patterns differ.
CD-MRFG addresses this narrow limitation while retaining the ASTGCN backbone. Its first contribution is a five-dimensional calendar-context encoding that separates periodic semantics from raw flow measurements. Its second contribution is a block-specific Softmax gate that converts the encoded context into weights for different Chebyshev orders. Its third contribution is an evaluation on PEMS03, PEMS04 and PEMS08 that combines reproduced baselines, module ablation, three-seed tests, gate diagnostics, complexity profiling, hyperparameter sensitivity and quantitative convergence analysis. The resulting claim is deliberately bounded: the method improves its direct ASTGCN baseline and exposes context-dependent order allocation, but it does not outperform every recent model.
The remainder of this paper is organized as follows. Section 2 reviews temporal, graph-based and context-aware forecasting methods and distinguishes CD-MRFG from related adaptive models. Section 3 defines the forecasting task and notation. Section 4 presents the model architecture and training objective. Section 5 reports the comparative, robustness, ablation, interpretation, efficiency, sensitivity and convergence results. Section 6 summarizes the findings and their limitations.

3. Problem Formulation

3.1. Traffic Network Graph

In the traffic forecasting task, the road network is represented as a weighted graph G = (V, E, A), where V is the set of sensor nodes, E denotes road connectivity and A is the adjacency matrix. PEMS03, PEMS04 and PEMS08 contain 358, 307 and 170 sensors, respectively, and all flow records are aggregated at 5 min intervals.
X t = x t 1 , , x t N T R N × C
Here, X t denotes the traffic state of all nodes at time step t, and C is the feature dimension. This study uses only traffic flow, so C = 1.

3.2. Multi-Step Traffic Forecasting Task

Given a sequence of L historical traffic-flow steps X t L + 1 : t , the graph G and the temporal context vector c t , the forecasting model learns a nonlinear mapping f_Theta to predict the next P time steps. Here, X t L + 1 : t belongs to R L × N × C , the predicted sequence belongs to R P × N × C , and both L and P are set to 12 in all experiments. Therefore, the model uses the previous 1 h to forecast the next 1 h.
X ^ t + 1 : t + P = f θ X t L + 1 : t , G , c t

3.3. Temporal Context Features

Traffic flow exhibits strong periodicity. The same sensor may follow different patterns across hours, weekdays and weekends. To represent this temporal context explicitly, a five-dimensional context vector is constructed at the forecasting origin:
c t = s h , q h , s d , q d , I weekend T R 5
s h = s i n 2 π h t 24 ,   q h = c o s 2 π h t 24
s d = s i n 2 π d t 7 ,   q d = c o s 2 π d t 7
where h t denotes the hour-of-day index in one day, d t denotes the day-of-week index in one week, and I w e e k e n d is the weekend indicator. The sine and cosine terms are used because hour and weekday are periodic variables.

3.4. Receptive Fields in Graph Convolution

Chebyshev graph convolution approximates a spectral graph filter with K polynomial orders. The k-th order can be interpreted as aggregating information from the k-hop neighborhood. The Chebyshev bases are defined recursively in Equations (6) and (7), where T k ( L ~ ) denotes the k-th Chebyshev polynomial and L ~ is the scaled graph Laplacian. Different orders, therefore, correspond to different spatial receptive fields, whose contributions need not be identical for short-term forecasting [22,23].
T 0 L ~ = I N ,   T 1 L ~ = L ~
T k L ~ = 2 L ~ T k 1 L ~ T k 2 L ~ ,   k 2

3.5. Evaluation Metrics

To evaluate multi-step traffic forecasting performance, this study uses mean absolute error (MAE), root mean square error (RMSE) and mean absolute percentage error (MAPE). MAE measures the average absolute deviation between predicted and observed traffic flows and provides a direct indication of overall forecasting accuracy. RMSE assigns larger penalties to large deviations and is, therefore, more sensitive to abrupt prediction errors around peak changes or congestion transitions. MAPE measures relative prediction error and is useful for comparing errors across different flow scales. Because traffic-flow records may contain zero or near-zero values, MAPE is calculated only on the valid set Ω, where the ground-truth value is larger than a small threshold. The three metrics are defined as follows:
M A E = 1 M i = 1 M y i y ^ i
R M S E = 1 M i = 1 M y i y ^ i 2
M A P E = 1 Ω i Ω y i y ^ i y i × 100 %
For all three metrics, lower values indicate better forecasting performance. MAPE is reported as a decimal fraction throughout the paper.

4. CD-MRFG: Context-Driven Dynamic Gating for ASTGCN

As shown in Figure 1, CD-MRFG consists of an ASTGCN spatio-temporal backbone, a context feature embedding module, a dynamic multi-order receptive-field gating module and a forecasting head. Historical traffic flow is processed by temporal attention, spatial attention, Chebyshev graph convolution and temporal convolution. The temporal-context branch explicitly displays the cyclic hour-of-day and day-of-week encodings rather than empirical traffic-flow profiles. The weekend indicator is represented as a binary variable, with zero for weekdays and one for weekend days. The resulting context representation generates sample-dependent Chebyshev-order weights for each ASTGCN block, allowing the model to adjust the spatial propagation range according to the temporal operating condition.
Figure 1. Overall architecture of CD-MRFG.

4.1. ASTGCN Spatio-Temporal Backbone

The ASTGCN backbone contains temporal attention, spatial attention, Chebyshev graph convolution, temporal convolution and residual connections [1]. Temporal attention selects informative historical steps, whereas spatial attention produces a block-specific matrix S b . Equation (11) gives the standard Chebyshev convolution. In CD-MRFG, the normalized spatial attention matrix is denoted by S b ^ , the order-specific propagation operator by M b , k , and the corresponding feature by Z b , k , as defined in Equations (14)–(16).
C h e b C o n v X = k = 0 K 1 T k L ~ X Θ k

4.2. Context Feature Embedding

The context embedding module maps c t to a sample-level hidden vector h c through a multilayer perceptron. This representation does not replace the traffic-flow input; instead, it provides conditional information for the gating network and captures differences in traffic propagation patterns under different temporal contexts.
h c = φ c t = M L P c t

4.3. Dynamic Multi-Order Receptive-Field Gating

For the b-th ASTGCN block, the gating network generates a K-dimensional weight vector g b from h c . After softmax normalization, the weights sum to one, allowing the model to select an appropriate spatial receptive field for each temporal context.
The motivation of the proposed gating mechanism is that different temporal contexts may require different spatial receptive fields. During peak hours, congestion may propagate to upstream and downstream sensors, so higher-order neighbourhoods may provide useful information. During off-peak periods, traffic states may be more locally stable, so lower-order neighbourhoods or self-information can be more reliable. Fixed Chebyshev-order weighting, therefore, restricts the modelling flexibility of graph convolution. By conditioning the order weights on temporal context, CD-MRFG provides a simple mechanism for representing context-dependent propagation ranges in a traffic network.
g b = S o f t m a x W g h c + b g , k = 0 K 1   g b , k = 1
M b , k = T k ( L ~ ) S b ^
Z b , k = M b , k T X b Θ b , k
Z b = k = 0 K 1 g b , k Z b , k
Figure 2 details this modelling mechanism. The five-dimensional temporal context vector is first encoded as a hidden representation. The gate then produces sample-dependent K-order weights for the b-th ASTGCN block. Each Chebyshev order corresponds to self-information, local neighbourhood propagation or more distant neighbourhood propagation. The weighted sum forms a gated graph-convolution output, which is combined with temporal convolution, residual connection and layer normalization to update spatio-temporal features.
Figure 2. Context-driven dynamic multi-order receptive-field gating.
Algorithm 1 summarizes the dynamic gated graph convolution in the b-th ASTGCN block. The procedure keeps the original spatial-attention graph-convolution framework but adds a context-conditioned modelling layer that controls how different Chebyshev basis functions contribute to the simulated traffic propagation process.
Algorithm 1. Context-conditioned dynamic multi-order receptive-field gating
Input: traffic-flow feature Xb, temporal context vector ct, Chebyshev polynomial set { T k ( L ~ ) } , and spatial attention matrix Sb
Output: gated graph-convolution feature Zb of the b-th ASTGCN block
1.hc ← MLPc(ct)
2.Ab ← MLPg,b(hc)
3.gb ← Softmax(Ab)
4. g b = [ g b , 0 , g b , 1 , , g b , K 1 ] , k = 0 K 1 g b , k = 1
5.Normalize Sb to obtain S b ^
6.Initialize Zb ← 0
7.for k = 0, 1, …, K − 1 do
8.  if k = 0 then
9.      Mb,kI
10.  else
11. M b , k T k ( L ~ ) S b ^
12.  end if
13.  Ub,k ← GraphConv(Xb, Mb,k, Θb,k)
14.  ZbZb + gb,k Ub,k
15.end for
16.Zb ← ReLU(Zb + βb)
17.return Zb

4.4. Training Objective

After the forecasting head produces the predicted sequence, CD-MRFG is optimized by minimizing the mean squared error between the predicted and observed traffic flows. Let y i and y ^ i denote the ground-truth and predicted values over all valid nodes, horizons and samples. The residual term and training loss are defined in Equations (17) and (18).
e i = y i y ^ i
L = 1 M i = 1 M e i 2

5. Experiments and Results

Experiments were conducted on PEMS03, PEMS04 and PEMS08. The evaluation addressed six questions: whether CD-MRFG improves its direct ASTGCN baseline, whether the result generalizes to a third sensor network, how it compares with recent strong baselines, whether the gains persist across random seeds, which components and hyperparameters control performance, and whether the added mechanism has acceptable computational and optimization costs. Unless otherwise stated, all reproduced models used the same chronological split, 12-step input, 12-step target and test metrics.

5.1. Datasets

PEMS03, PEMS04 and PEMS08 are public California freeway detector datasets sampled every 5 min. As shown in Table 2, PEMS03 contains 358 sensors and 26,208 time steps from 1 September to 30 November 2018. PEMS04 contains 307 sensors and 16,992 time steps from 1 January to 28 February 2018, and PEMS08 contains 170 sensors and 17,856 time steps from 1 July to 31 August 2016. Each dataset was divided chronologically into training, validation and test sets using a 6:2:2 ratio. Twelve historical observations were used to forecast the following twelve observations.
Table 2. Dataset statistics.

5.2. Experimental Settings and Metrics

Experiments were run on a workstation with a 13th Gen Intel Core i7-13700K CPU and an NVIDIA GeForce RTX 4090D GPU. Architecture-specific objectives and validation-selected settings followed the corresponding implementations, while the data split, input length, forecast horizon and evaluation code were fixed. As shown in Table 3, ASTGCN and CD-MRFG used Adam, mean squared error, a learning rate of 0.001 and a batch size of 32. Their common configuration contained two spatio-temporal blocks, Chebyshev order K = 3, 64 graph and temporal filters, and a CD-MRFG context hidden dimension of 64. The controlled benchmark runs of CD-MRFG achieved an overall MAE and RMSE of 16.98 and 28.59 on PEMS03, 20.82 and 32.77 on PEMS04, and 17.24 and 26.84 on PEMS08. Repeated-seed statistics are reported separately in Section 5.4.
Table 3. Experimental settings.
Following common practice in traffic forecasting, mean absolute error (MAE), root mean square error (RMSE) and masked mean absolute percentage error (MAPE) are used as metrics. Their formal definitions are given in Equations (8)–(10). Let Y and Ŷ denote the ground-truth and predicted traffic flows, N the number of sensors and T the prediction horizon. Because traffic flow may contain zero or near-zero values, MAPE is calculated only on the valid set Ω, where the ground-truth value is larger than ε. In this paper, ε is set to 1 × 10−5. A lower MAE and RMSE indicate smaller forecasting error. RMSE is more sensitive to large errors, whereas MAPE measures relative error.

5.3. Comparison with Baseline Methods

The comparison includes HA, LSTM, GRU, DCRNN, STGCN, ASTGCN, AGCRN and STSGCN. All values in Table 4 were reproduced with the common chronological split and evaluation code. Model-specific architecture and optimization settings followed their implementations and were selected using validation data. ASTGCN is the primary controlled baseline because CD-MRFG modifies its graph-convolution stage. STID and MTGNN are reported separately in Table 5 as recent strong baselines on PEMS04 and PEMS08.
Table 4. Forecasting performance of CD-MRFG and reproduced baselines on PEMS03, PEMS04 and PEMS08 datasets.
Table 5. Three-seed robustness and paired significance tests for ASTGCN and CD-MRFG.
Errors increased with the prediction horizon for the learned baselines on all three datasets. On PEMS03, CD-MRFG achieved an overall MAE and RMSE of 16.98 and 28.59. It improved DCRNN by 5.67% in MAE and 1.68% in RMSE, STGCN by 15.48% and 12.00%, and STSGCN by 4.34% and 2.46%, respectively. AGCRN remained more accurate, with an overall MAE and RMSE of 15.51 and 27.17. The third dataset, therefore, supports generalization beyond PEMS04 and PEMS08 while preserving the bounded accuracy claim.
On PEMS04, CD-MRFG reduced overall MAE and RMSE by 6.13% and 4.25% relative to DCRNN, and by 20.93% and 19.32% relative to STGCN. On PEMS08, it reduced MAE by 0.52% relative to DCRNN, whereas its RMSE was 0.60% higher. AGCRN produced lower overall errors than CD-MRFG on PEMS04 and PEMS08. These mixed results show that the proposed gate improves an ASTGCN-style model but does not dominate adaptive graph learning.
The direct ASTGCN comparison was consistent across datasets. CD-MRFG reduced overall MAE and RMSE from 18.66 and 31.05 to 16.98 and 28.59 on PEMS03, from 22.79 and 35.02 to 20.82 and 32.77 on PEMS04, and from 18.88 and 28.83 to 17.24 and 26.84 on PEMS08. The corresponding MAE reductions were 9.00%, 8.64% and 8.69%. This controlled comparison supports the specific claim that explicit calendar context and order gating improve the ASTGCN backbone.
Table 6 provides a stricter comparison with STID and MTGNN. Both recent baselines achieved lower overall errors than CD-MRFG on PEMS04 and PEMS08. The proposed method is, therefore, evaluated for its controlled gain over ASTGCN, explicit order-level diagnostics and modest architectural change, not for state-of-the-art accuracy.
Table 6. Overall comparison with reproduced recent strong baselines on PEMS04 and PEMS08.
MTGNN and STID benefit from learned graph structure or strong identity embeddings and, therefore, define an important accuracy boundary. Their advantage does not invalidate the controlled ASTGCN comparison, but it prevents a universal superiority claim.

5.4. Robustness Across Random Seeds

ASTGCN and CD-MRFG were independently trained with seeds 1, 42 and 2026 on PEMS04 and PEMS08. Table 5 reports mean ± standard deviation over the three runs. Paired two-sided t-tests compare the two models using matched seeds. Given the small sample size, the p-values are interpreted together with the paired differences rather than as standalone proof.
CD-MRFG reduced the mean MAE for every matched seed. The MAE difference was significant on PEMS04 (p = 0.028) and PEMS08 (p = 0.042). The mean RMSE was also lower, but the paired tests did not cross the 0.05 threshold on PEMS04 (p = 0.181) or PEMS08 (p = 0.057). MAPE differences were not significant. These results support a reproducible MAE gain while placing an explicit statistical boundary on the RMSE and MAPE claims.

5.5. Ablation Study

Four variants are evaluated on PEMS04 and PEMS08 to isolate the effect of each component. M1 is the original ASTGCN baseline. M2 adds fixed multi-order graph-convolution gating to ASTGCN. M3 uses only the context feature embedding module. M4, namely CD-MRFG, combines context embedding with dynamic multi-order receptive-field gating. Table 7 reports the average 12-step results.
Table 7. Average 12-step results of the ablation variants on PEMS04 and PEMS08.
  • M1 ASTGCN. This baseline measures the forecasting ability of the original spatio-temporal attention graph convolution without temporal context or multi-order gating. Its average MAE is 22.79 and RMSE is 35.02 on PEMS04, while its average MAE is 18.88 and RMSE is 28.83 on PEMS08. These results show that ASTGCN captures basic spatio-temporal dependencies, but its input remains dominated by historical flow values and does not explicitly encode temporal context.
  • M2 fixed multi-order graph-convolution gating. M2 does not provide a neutral modification to the ASTGCN baseline. On PEMS04, its MAE remains nearly unchanged, decreasing from 22.79 to 22.78, whereas its RMSE increases from 35.02 to 36.45, corresponding to a relative deterioration of 4.08%. Because RMSE assigns greater weight to large residuals, this divergence indicates that fixed order weighting increases the magnitude of difficult forecasting errors even when the average absolute error remains stable. A fixed gate applies the same receptive-field mixture to all samples and cannot distinguish traffic states dominated by local variation from states involving broader spatial propagation. This mismatch can cause excessive neighbourhood aggregation in local regimes or insufficient higher-order propagation during spatially extended congestion. On PEMS08, M2 produces only marginal reductions in MAE and RMSE, from 18.88 and 28.83 to 18.85 and 28.56, respectively, while MAPE increases from 0.12 to 0.13. The opposite RMSE changes across datasets indicate that a fixed receptive-field mixture interacts with dataset-specific topology and traffic-regime composition rather than providing a robust improvement. In contrast, M4 conditions the order weights on temporal context and reduces the PEMS04 RMSE from 36.45 to 32.77 and the PEMS08 RMSE from 28.56 to 26.84 relative to M2. These results indicate that the effectiveness of multi-order gating depends on sample-specific contextual modulation rather than on the introduction of additional order weights alone.
  • M3 context feature embedding. This variant evaluates the contribution of temporal context itself. M3 reduces MAE to 21.34 and RMSE to 33.81 on PEMS04 and reduces MAE to 17.83 and RMSE to 27.25 on PEMS08, clearly outperforming M1 on both datasets. Hour, weekday and weekend indicators, therefore, provide useful periodic semantics that help distinguish morning and evening peaks, off-peak periods and weekend traffic.
  • M4 CD-MRFG. The full model combines context embedding with dynamic multi-order receptive-field gating and achieves the best results among the four variants. It obtains an MAE of 20.82, RMSE of 32.77 and MAPE of 0.14 on PEMS04, and an MAE of 17.24 and RMSE of 26.84 on PEMS08. Its further improvement over M3 indicates that context is useful not only as an auxiliary input representation but also as a condition for selecting graph-convolution orders.

5.6. Gate Weight Analysis Under Different Traffic Periods

Gate weights were extracted from the PEMS04 and PEMS08 test sets to determine whether order allocation changed with temporal context. Table 8 reports the mean weights assigned to orders 0, 1 and 2 in the two ASTGCN blocks for four traffic periods. Each row sums to one because the gate uses a Softmax output.
Table 8. Mean dynamic gate weights under different traffic periods.
The distributions varied across traffic periods and blocks. On PEMS04, block 1 emphasized order 2 during the morning peak and daytime off-peak periods, whereas block 2 emphasized order 1 in the morning peak and order 0 in the evening peak and at night. On PEMS08, block 1 mainly balanced orders 0 and 1, while block 2 assigned larger weights to order 2 in the three daytime periods. The gate, therefore, did not collapse to a single fixed mixture.
These weights are model-internal allocation coefficients, not measurements of physical traffic propagation or causal influence. Their variation supports the narrower conclusion that CD-MRFG uses different polynomial graph components under different calendar contexts. Physical interpretation would require external traffic-state evidence or controlled interventions.

5.7. Model Interpretation and Applicability Boundary

CD-MRFG is designed for short-term forecasting on fixed sensor graphs with regular daily and weekly regimes. The gate converts calendar context into a convex combination of self, first-order and higher-order graph features. This construction makes the model’s receptive-field preference inspectable without replacing the ASTGCN backbone.
The applicability boundary is equally important. The five context variables do not represent holidays, incidents, weather, roadworks or special events. The method may, therefore, be less reliable when non-recurring disturbances dominate the traffic state. It also retains a predefined graph and cannot recover missing or evolving connectivity as directly as adaptive graph-learning models.
AGCRN, STID and MTGNN achieved lower absolute errors in several comparisons. CD-MRFG should, thus, be used when the objective is a controlled ASTGCN improvement with explicit order-level diagnostics and modest structural change, rather than when forecasting accuracy alone determines model choice.

5.8. Complexity and Computational Overhead

Because CD-MRFG adds only a context encoder and a dynamic receptive-field gate to ASTGCN, its overhead mainly comes from lightweight multilayer perceptrons and scalar order-wise weighting. Table 9 reports profiling under a batch size of 32 on CUDA. On PEMS04, the parameter count increases from 0.450 M to 0.459 M, and latency increases from 24.06 ms to 27.86 ms. On PEMS08, the parameter count increases from 0.179 M to 0.189 M, and latency increases from 17.94 ms to 19.53 ms. The measured PEMS08 checkpoint sizes are 0.6966 MB for ASTGCN and 0.7342 MB for CD-MRFG. The measured values indicate limited model-size growth and moderate inference overhead.
Table 9. Parameter scale and computational profiling results.

5.9. Hyperparameter Sensitivity

A one-factor-at-a-time analysis was conducted on PEMS04 with seed 42. Chebyshev order K, context hidden dimension and the number of spatio-temporal blocks were varied while the remaining settings were fixed. The analysis was performed after the main benchmark and was not used to retrospectively select the reported test configuration. Table 10 reports PEMS04 hyperparameter sensitivity results and Figure 3 demonstrates PEMS04 sensitivity to the varied parameters.
Table 10. PEMS04 hyperparameter sensitivity results.
Figure 3. PEMS04 sensitivity to Chebyshev order, context hidden dimension and number of spatio-temporal blocks. Orange markers denote the fixed benchmark settings.
Increasing K from 2 to 4 progressively reduced MAE and RMSE, but K = 4 increased training time by 11.1% relative to K = 3. A hidden dimension of 32 was more accurate and smaller than the default 64, whereas 128 degraded both errors, indicating that a larger context encoder is not automatically beneficial. Three blocks achieved the lowest errors but increased parameters and training time by about 52% and 50% relative to two blocks. The main setting, therefore, represents a fixed benchmark configuration rather than the best point from post hoc tuning.

5.10. Quantitative Convergence Analysis

Table 11 reports quantitative convergence results from the completed three-seed logs. E90 and E95 denote the first epochs that reached 90% and 95% of the final achieved validation-loss reduction. In Figure 4, normalized area under the validation curve (AUC) summarizes the full optimization trajectory, with a smaller value indicating faster loss reduction. Late-stage variability is measured by the coefficient of variation over the final ten epochs.
Table 11. Quantitative convergence results over three random seeds.
Figure 4. Normalized validation-loss trajectories across three seeds on (a) PEMS04 and (b) PEMS08. Shaded regions denote mean ± one standard deviation.
Both models reached E90 within approximately five epochs. CD-MRFG had a lower normalized AUC on both datasets. The paired difference was significant on PEMS08 (p = 0.031) and marginal on PEMS04 (p = 0.054). On PEMS08, CD-MRFG also reduced late-stage variability from 3.18% to 0.84%. These optimization benefits required additional time: the mean training duration increased by 4.58 min on PEMS04 and 4.16 min on PEMS08. The results support a more efficient loss trajectory, particularly on PEMS08, rather than a universal claim of faster threshold attainment.

5.11. Forecasting Visualization

In Figure 5 and Figure 6, representative PEMS04 and PEMS08 test sequences are plotted with ground-truth flow, ASTGCN predictions and CD-MRFG predictions. The curves provide qualitative evidence about turning-point response and trend continuity. PEMS03 is evaluated quantitatively in Table 4 but is not used to select an additional visualization.
Figure 5. Forecasting visualization on PEMS04.
Figure 6. Forecasting visualization on PEMS08.
On PEMS04, the selected sequence contains evident morning-like growth, a high-flow plateau, subsequent decline and recovery. These stages are useful for evaluating multi-step forecasting because errors easily accumulate around turning points. The ASTGCN baseline captures the broad temporal trend, but its predicted curve tends to lag when the flow changes quickly. CD-MRFG gives a smoother yet more responsive trajectory, especially around the falling segment after the local peak and the later recovery interval.
The PEMS08 visualization provides a second case with a different sensor scale and temporal distribution. Although the overall profile differs from PEMS04, the same phenomenon can be observed: baseline predictions are delayed in several peak-decline and trough-recovery intervals, whereas CD-MRFG remains closer to the ground-truth direction of change. This cross-dataset consistency supports the argument that the proposed context-conditioned gate improves the dynamic selection of spatial receptive fields instead of overfitting one particular dataset.
The visual results complement the aggregate errors in Table 4. They show that CD-MRFG follows several rise, decline and recovery segments more closely than ASTGCN. These examples are illustrative rather than statistical evidence and should be interpreted together with the full test-set metrics and repeated-seed analysis.

6. Conclusions

This paper presented CD-MRFG, a context-gated ASTGCN extension that maps calendar information to Chebyshev-order weights. Controlled experiments on PEMS03, PEMS04 and PEMS08 consistently improved the direct ASTGCN baseline, and three-seed tests confirmed significant mean MAE reductions on PEMS04 and PEMS08. Ablation identified temporal context as the main contributor, while gate diagnostics showed non-constant order allocation across periods and blocks. Sensitivity and convergence analyses further exposed the effects of model scale and the associated optimization cost. CD-MRFG did not outperform AGCRN, STID or MTGNN in every setting, and its gate weights are not causal measurements of traffic propagation. The method is, therefore, most appropriate as a lightweight and inspectable improvement to ASTGCN on fixed sensor graphs with recurring calendar regimes.

Author Contributions

Methodology, Y.Z.; software, J.L.; validation, Z.Y., Z.D. and Y.K.; data curation, Z.D.; writing—original draft preparation, J.L.; writing—review and editing, Y.Z.; visualization, Y.Z.; supervision, Y.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by National Natural Science Foundation of China, grant number 52302408.

Data Availability Statement

The PEMS03, PEMS04 and PEMS08 datasets used in this study are public traffic forecasting benchmarks. The processed splits follow the chronological 6:2:2 protocol described in Section 5.1.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ASTGCNAttention-based Spatial–Temporal Graph Convolutional Network
CD-MRFGContext-Driven Dynamic Multi-Order Receptive-Field Gated ASTGCN
PEMS04Performance Measurement System Dataset 04
PEMS08Performance Measurement System Dataset 08
HAHistorical Average
LSTMLong Short-Term Memory
GRUGated Recurrent Unit
GCNGraph Convolutional Network
STGCNSpatio-Temporal Graph Convolutional Network
DCRNNDiffusion Convolutional Recurrent Neural Network
AGCRNAdaptive Graph Convolutional Recurrent Network
STSGCNSpatial–Temporal Synchronous Graph Convolutional Network
STIDSpatial–Temporal Identity
GMANGraph Multi-Attention Network
STDNSpatial–Temporal Dynamic Network
T-GCNTemporal Graph Convolutional Network
PEMS03Performance Measurement System Dataset 03
MTGNNMultivariate Time-Series Graph Neural Network
PDFormerPropagation Delay-Aware Dynamic Long-Range Transformer
STAEformerSpatio-Temporal Adaptive Embedding Transformer
AUCArea Under the Curve
MRA-BGCNMulti-Range Attentive Bicomponent Graph Convolutional Network

References

  1. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv 2014, arXiv:1412.3555. [Google Scholar] [CrossRef] [Scilit]
  3. Guo, S.; Lin, Y.; Feng, N.; Song, C.; Wan, H. Attention Based Spatial-Temporal Graph Convolutional Networks for Traffic Flow Forecasting. Proc. AAAI Conf. Artif. Intell. 2019, 33, 922–929. [Google Scholar] [CrossRef] [Scilit]
  4. Yu, B.; Yin, H.; Zhu, Z. Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden, 13–19 July 2018; AAAI Press: Washington, DC, USA, 2018; pp. 3634–3640. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Li, Y.; Yu, R.; Shahabi, C.; Liu, Y. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018; ICLR: Appleton, WI, USA, 2018; pp. 1–16. [Google Scholar] [CrossRef] [Scilit]
  6. Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Zhang, C. Graph WaveNet for Deep Spatio-Temporal Graph Modeling. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, Macau, China, 10–16 August 2019; AAAI Press: Washington, DC, USA, 2019; pp. 1907–1913. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Bai, L.; Yao, L.; Li, C.; Wang, X.; Wang, C. Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting. Adv. Neural Inf. Process. Syst. 2020, 33, 17804–17815. [Google Scholar] [CrossRef] [Scilit]
  8. Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Chang, X.; Zhang, C. Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Virtual Event, CA, USA, 6–10 July 2020; ACM: New York, NY, USA, 2020; pp. 753–763. [Google Scholar] [CrossRef] [Scilit]
  9. Li, W.; Chen, J.; Zhang, Y.; Sun, R.; Xia, S.; Pan, Z.; Luo, J. MSGFormer: Revolutionizing Traffic Flow Prediction with Multiscale and Gated Transformer Architecture. IEEE Internet Things J. 2025, 12, 2014–2025. [Google Scholar] [CrossRef] [Scilit]
  10. Zheng, C.P.; Fan, X.L.; Wang, C.; Qi, J.Z. GMAN: A Graph Multi-Attention Network for Traffic Prediction. Proc. AAAI Conf. Artif. Intell. 2020, 34, 1234–1241. [Google Scholar] [CrossRef] [Scilit]
  11. Jiang, J.; Han, C.; Zhao, W.X.; Wang, J. PDFormer: Propagation Delay-Aware Dynamic Long-Range Transformer for Traffic Flow Prediction. Proc. AAAI Conf. Artif. Intell. 2023, 37, 4365–4373. [Google Scholar] [CrossRef] [Scilit]
  12. Liu, H.; Dong, Z.; Jiang, R.; Deng, J.; Deng, J.; Chen, Q.; Song, X. Spatio-Temporal Adaptive Embedding Makes Vanilla Transformer SOTA for Traffic Forecasting. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, Birmingham, UK, 21–25 October 2023; ACM: New York, NY, USA, 2023; pp. 4125–4129. [Google Scholar] [CrossRef] [Scilit]
  13. Song, C.; Lin, Y.; Guo, S.; Wan, H. Spatial-Temporal Synchronous Graph Convolutional Networks: A New Framework for Spatial-Temporal Network Data Forecasting. Proc. AAAI Conf. Artif. Intell. 2020, 34, 914–921. [Google Scholar] [CrossRef] [Scilit]
  14. Guo, K.; Hu, Y.; Sun, Y.; Qian, Z.S.; Gao, J.; Yin, B. An Optimized Temporal-Spatial Gated Graph Convolution Network for Traffic Forecasting. IEEE Intell. Transp. Syst. Mag. 2022, 14, 153–162. [Google Scholar] [CrossRef]
  15. Zhao, L.; Song, Y.; Zhang, C.; Liu, Y.; Wang, P.; Lin, T.; Deng, M.; Li, H. T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction. IEEE Trans. Intell. Transp. 2020, 21, 3848–3858. [Google Scholar] [CrossRef] [Scilit]
  16. Zhao, Y.; Lin, Y.; Wen, H.; Wei, T.; Jin, X.; Wan, H. Spatial-Temporal Position-Aware Graph Convolution Networks for Traffic Flow Forecasting. IEEE Trans. Intell. Transp. Syst. 2023, 24, 8650–8666. [Google Scholar] [CrossRef] [Scilit]
  17. Chen, W.; Chen, L.; Xie, Y.; Cao, W.; Gao, Y.; Feng, X. Multi-Range Attentive Bicomponent Graph Convolutional Network for Traffic Forecasting. Proc. AAAI Conf. Artif. Intell. 2020, 34, 3529–3536. [Google Scholar] [CrossRef] [Scilit]
  18. Shao, Z.; Zhang, Z.; Wang, F.; Wei, W.; Xu, Y. Spatial-Temporal Identity: A Simple yet Effective Baseline for Multivariate Time Series Forecasting. arXiv 2022, arXiv:2208.05233. [Google Scholar] [CrossRef] [Scilit]
  19. Cai, L.; Janowicz, K.; Mai, G.; Yan, B.; Zhu, R. Traffic Transformer: Capturing the Continuity and Periodicity of Time Series for Traffic Forecasting. Trans. GIS 2020, 24, 736–755. [Google Scholar] [CrossRef] [Scilit]
  20. Zhang, J.B.; Zheng, Y.; Qi, D.K. Deep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction. Proc. AAAI Conf. Artif. Intell. 2017, 31, 1655–1661. [Google Scholar] [CrossRef] [Scilit]
  21. Yao, H.X.; Tang, X.F.; Wei, H.; Zheng, G.; Li, Z. Revisiting Spatial-Temporal Similarity: A Deep Learning Framework for Traffic Prediction. Proc. AAAI Conf. Artif. Intell. 2019, 33, 5668–5675. [Google Scholar] [CrossRef] [Scilit]
  22. Defferrard, M.; Bresson, X.; Vandergheynst, P. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In Advances in Neural Information Processing Systems; NeurIPS: Barcelona, Spain, 2016; pp. 3844–3852. [Google Scholar] [CrossRef] [Scilit]
  23. Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the International Conference on Learning Representations; ICLR: Toulon, France, 2017; pp. 1–14. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.