A Dual-Attentional Gated Residual Framework for Robust Travel Time Prediction
Abstract
1. Introduction
- Physics-Aware Structural Prior: This tackles the lack of spatial observability by explicitly encoding road segment connectivity into a directed Line Graph. This provides a hard structural prior that inherently respects physical traffic constraints without requiring massive training data to infer spatial topologies.
- Dynamic Context Modulation: This mitigates environmental volatility by dynamically injecting external contexts (e.g., peak-hour time features and meteorological conditions) via Feature-wise Linear Modulation (FiLM).
- Multi-Head Attention Residual Fusion: This overcomes the gradient vanishing problem inherent in ultra-long trajectories through a gated residual mechanism that dynamically balances global Bidirectional GRU sequence representations with an environment-guided Multi-Head Attention module, effectively capturing extreme localized congestions.
- Exceptional Cold-Start Generalization: Evaluated on a highly constrained, stratified real-world dataset comprising merely 2000 trajectories, DAGRN achieves a state-of-the-art RMSE of 415.485 s and an R2 of 0.848, unequivocally demonstrating its superior deployability in data-scarce ITS environments.
2. Research Methodology
2.1. Data Preprocessing and Graph Construction
2.1.1. Line Graph Transformation
2.1.2. Trajectory Processing and Stratified Evaluation
- Temporal Stratification (Figure 3b): Captures volatility during peak hours (7–9 a.m., 17–19 p.m.).
- Spatial Stratification (Figure 3c): Specifically targets the spatial complexity by explicitly partitioning the long-distance trips (>20 links). According to our dataset distribution, this specific subset accounts for 57.6% of the total trajectories, ensuring the model effectively captures extended spatial dependencies without being biased by fragmented short trips.
- Traffic State Stratification (Figure 3d): Isolates congested flows where average speeds drop below 20 km/h. This cut-off is empirically derived and rigorously defines severe congestion scenarios, which constitute the most critical 9.7% of the overall speed distribution.
2.2. The DAGRN Architecture
2.2.1. Gated Graph Dynamics with FiLM Modulation
2.2.2. Bidirectional Sequence Backbone
2.2.3. Multi-Head Attention Residual Fusion
2.3. Optimization and Training Strategy
3. Experiments and Analysis
3.1. Experimental Setup
3.1.1. Dataset and Evaluation Metrics
3.1.2. Data Stratification for Robustness Analysis
- Sequence and Temporal Models. We evaluate fundamental recurrent architectures, including LSTM and GRU. Recent works [35] have highlighted their continued relevance for handling missing data in sparse environments. We also include Temporal Convolutional Networks (TCNs) [36] and attention-enhanced variants (Attn-GRU).
- Spatio-Temporal GNNs. We include advanced graph-based architectures that fuse spatial convolution with recurrent units, namely GCN-GRU, TAGCN-GRU, GAT-LSTM [37], and Graph WaveNet. Additionally, Additionally, we evaluate Adaptive-GGNN as a representative baseline for adaptive graph recurrent networks [38].
3.2. Comparative Analysis
3.2.1. Performance Analysis on Sparse Data
3.2.2. Robustness in Challenging Scenarios
- (a)
- Long-Distance Trips (>20 links): Longer trajectories exponentially amplify cumulative prediction errors. DAGRN mitigates this error accumulation via its Gated Residual Fusion mechanism, delivering consistently low RMSE and actively avoiding the sharp error spikes observed in baselines.
- (b)
- Peak-Hour Traffic: During peak hours (07:00–09:00 and 17:00–19:00), traffic patterns exhibit heightened volatility. DAGRN’s Multi-Head Attention effectively captures these localized fluctuations, maintaining stable Peak-Hour RMSE despite limited training samples.
- (c)
- Congested Regimes: In severely congested conditions (defined as the top 10% slowest trips), prediction becomes highly sensitive to minor speed variances. DAGRN sustains highly competitive Congested-RMSE and MAPE metrics, yielding fewer extreme outliers than most baselines, thereby indicating strong robustness against traffic degradation.
3.3. Ablation Study
- Impact of Data Stratification (w/o Stratification): Training the model on a randomly sampled dataset without stratification exposes a dangerous long-tail bias, with the RMSE significantly increasing from 415.49 s to 429.37 s. This degradation is most pronounced in the severely congested regime, where the Congested-RMSE drastically spikes from 377.31 s to 431.18 s. This explicitly validates that our spatial stratification strategy effectively prevents the model from overwhelmingly overfitting to abundant, trivial short trips, guaranteeing robust generalization on complex, volatile trajectories.
- Impact of Topological Structure (w/o GCN): Removing the GCN module causes a significant performance drop, with the RMSE increasing to 444.17 s and R2 dropping to 0.826. This unequivocally confirms that the Line Graph topology serves as a foundational regularizer. As noted by Ji et al. [45], incorporating physical connectivity constraints is essential when training data is insufficient to learn the underlying manifold.
- Impact of Dynamic Context (w/o FiLM): The removal of FiLM modulation increases the RMSE to 447.67 s. This confirms that static topological embeddings alone are insufficient; the graph structure must be dynamically modulated by environmental contexts to prevent systematic bias under varying weather or temporal conditions.
- Impact of Attention and Fusion (w/o Attn, w/o Gate): The gating mechanism is structurally indispensable, as its removal (w/o Gate) causes the overall RMSE to regress to 454.289 s. More critically, the complete removal of the Multi-Head Attention module (w/o Attn) precipitates the most severe performance collapse across the entire ablation study, with the overall RMSE skyrocketing to 521.495 s. Furthermore, without the environment-guided local attention, the model loses its ability to capture intersection delays, causing the Congested-RMSE to surge to 429.464 s.
3.4. Parameter Sensitivity Analysis
- Impact of Model Capacity: Increasing the hidden dimension from 32 to 128 yields a consistent performance boost, suggesting the model benefits from a larger representational capacity to capture complex topological dependencies. However, the improvement plateaus, indicating that a dimension of 64 provides a cost-effective balance for practical deployment. This observation is consistent with Di et al. [46], who suggested that compact representations are often more robust for traffic state estimation under data constraints.
- Impact of Regularization: The dropout rate exhibits a distinct preference for lower values. Contrary to the common assumption that high dropout is essential for small datasets, a rate of 0.2 achieved optimal performance, whereas increasing it to 0.5 or 0.7 significantly deteriorated the RMSE. This indicates that the structural priors introduced by our Line Graph and FiLM modules actively function as strong regularizers; consequently, aggressive dropout merely disrupts the learning of subtle traffic patterns.
- Impact of Optimization Settings: A learning rate of 0.001 emerges as the most stable optimization parameter. Notably, a smaller batch size of 32 significantly outperforms larger batches (e.g., 128). This phenomenon is highly characteristic of data-scarce regimes, where smaller batches provide more frequent weight updates and introduce stochastic noise that assists the optimizer in escaping sharp local minima.
3.5. Computational Complexity Analysis
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| DAGRN | Dual-Attention Gated Residual Network |
| TTP | travel time prediction |
| ITS | Intelligent Transportation Systems |
| HMM | Hidden Markov Model |
| GGNN | Gated Graph Neural Network |
| FiLM | Feature-wise Linear Modulation |
| Bi-GRU | Bidirectional Gated Recurrent Unit |
| LLM | Large Language Model |
References
- Billings, D.; Yang, J.-S. Application of the ARIMA models to urban roadway travel time prediction—A case study. In Proceedings of the 2006 IEEE International Conference on Systems, Man and Cybernetics, Taipei, Taiwan, 8–11 October 2006; IEEE: Piscataway, NJ, USA, 2006; pp. 2529–2534. [Google Scholar] [CrossRef] [Scilit]
- Guin, A. Travel time prediction using a seasonal autoregressive integrated moving average time series model. In Proceedings of the 2006 IEEE Intelligent Transportation Systems Conference, Toronto, ON, Canada, 17–20 September 2006; IEEE: Piscataway, NJ, USA, 2006; pp. 493–498. [Google Scholar] [CrossRef] [Scilit]
- Yin, X.; Wu, G.; Wei, J.; Shen, Y.; Qi, H.; Yin, B. Deep learning on traffic prediction: Methods, analysis, and future directions. IEEE Trans. Intell. Transp. Syst. 2022, 23, 4927–4943. [Google Scholar] [CrossRef] [Scilit]
- Wu, C.-H.; Ho, J.-M.; Lee, D.T. Travel-time prediction with support vector regression. IEEE Trans. Intell. Transp. Syst. 2004, 5, 276–281. [Google Scholar] [CrossRef] [Scilit]
- Fei, X.; Lu, C.-C.; Liu, K. A Bayesian dynamic linear model approach for real-time short-term freeway travel time prediction. Transp. Res. Part C Emerg. Technol. 2011, 19, 1306–1318. [Google Scholar] [CrossRef] [Scilit]
- Jiang, R.; Yin, D.; Wang, Z.; Wang, Y.; Deng, J.; Liu, H.; Cai, Z.; Deng, J.; Song, X.; Shibasaki, R. DL-Traff: Survey and benchmark of deep learning models for urban traffic prediction. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management (CIKM), Gold Coast, QLD, Australia, 1–5 November 2021; ACM: New York, NY, USA, 2021; pp. 4515–4525. [Google Scholar] [CrossRef] [Scilit]
- Tedjopurnomo, D.A.; Bao, Z.; Zheng, B.; Choudhury, F.M.; Qin, A.K. A survey on modern deep neural network for traffic prediction: Trends, methods and challenges. IEEE Trans. Knowl. Data Eng. 2022, 34, 1544–1561. [Google Scholar] [CrossRef] [Scilit]
- Duan, Y.; Lv, Y.; Wang, F.-Y. Travel time prediction with LSTM neural network. In Proceedings of the 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC), Rio de Janeiro, Brazil, 1–4 November 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 1053–1058. [Google Scholar] [CrossRef] [Scilit]
- Abdellah, A.R.; Abdelmoaty, A.; Ateya, A.A.; Abd El-Latif, A.A.; Muthanna, A.; Koucheryavy, A. Accurate V2X traffic prediction with deep learning architectures. Front. Artif. Intell. 2025, 8, 1565287. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, Y.; Yu, R.; Shahabi, C.; Liu, Y. Diffusion convolutional recurrent neural network: Data-driven traffic forecasting. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar] [CrossRef] [Scilit]
- Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Zhang, C. Graph WaveNet for deep spatial-temporal graph modeling. arXiv 2019, arXiv:1906.00121. [Google Scholar] [CrossRef] [Scilit]
- Lan, S.; Ma, Y.; Huang, W.; Wang, W.; Yang, H.; Li, P. DSTAGNN: Dynamic spatial-temporal aware graph neural network for traffic flow forecasting. In Proceedings of the 39th International Conference on Machine Learning, Baltimore, MD, USA, 17–23 July 2022; Available online: https://proceedings.mlr.press/v162/lan22a (accessed on 12 January 2026).
- Jiang, J.; Han, C.; Zhao, W.X.; Wang, J. PDFormer: Propagation delay-aware dynamic long-range transformer for traffic flow prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023. [Google Scholar] [CrossRef] [Scilit]
- Li, L.; Wang, H.; Zhang, W.; Coster, A. STG-Mamba: Spatial-temporal graph learning via selective state space model. arXiv 2024, arXiv:2403.12418. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Yang, Y.; Wu, X.; Wang, H. K-shape-based spatio-temporal graph attention network for traffic flow prediction. In Proceedings of the Fourth International Conference on Intelligent Traffic Systems and Smart City (ITSSC), Nanjing, China, 24–26 May 2024. [Google Scholar] [CrossRef] [Scilit]
- Kong, L.; Yang, H.; Li, W.; Zhang, Y.; Guan, J.; Zhou, S. Traffexplainer: A framework toward GNN-based interpretable traffic prediction. IEEE Trans. Artif. Intell. 2025, 6, 559–573. [Google Scholar] [CrossRef] [Scilit]
- Lu, S.; Chen, H.; Teng, Y. Multi-Scale Non-Local Spatio-Temporal Information Fusion Networks for Multi-Step Traffic Flow Forecasting. ISPRS Int. J. Geo-Inf. 2024, 13, 71. [Google Scholar] [CrossRef] [Scilit]
- Zouari, M.; Baklouti, N.; Sanchez-Medina, J.; Kammoun, H.M.; Ayed, M.B.; Alimi, A.M. PSO-based adaptive hierarchical interval type-2 fuzzy knowledge representation system (PSO-AHIT2FKRS) for travel route guidance. IEEE Trans. Intell. Transp. Syst. 2022, 23, 804–818. [Google Scholar] [CrossRef] [Scilit]
- Akopov, A.S.; Beklaryan, L.A. Evolutionary synthesis of high-capacity reconfigurable multilayer road networks using a multiagent hybrid clustering-assisted genetic algorithm. IEEE Access 2025, 13, 53448–53474. [Google Scholar] [CrossRef] [Scilit]
- Luo, Q.; Wang, H.; Yang, J.; Zang, X.; Chen, X.; Postolache, O. Study on a multi-factor lane-changing risk resilience assessment model based on genetic algorithm and fault tree analysis. Accid. Anal. Prev. 2026, 230, 108443. [Google Scholar] [CrossRef] [Scilit]
- Luo, Q.; Lu, X.; Zang, Z.; Gong, H.; Guo, X.; Chen, X. A Real-Time Early Warning Framework for Multi-Dimensional Driving Risk of Heavy-Duty Trucks Using Trajectory Data. Systems 2026, 14, 204. [Google Scholar] [CrossRef] [Scilit]
- Fan, Y.; Xu, J.; Zhou, R.; Li, J.; Zheng, K.; Chen, L.; Liu, C. MetaER-TTE: An adaptive meta-learning model for en route travel time estimation. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI), Vienna, Austria, 23–29 July 2022; Available online: https://zheng-kai.com/paper/ijcai_2022_fan.pdf (accessed on 12 January 2026).
- Liu, M.; Huang, H.; Feng, H.; Sun, L.; Du, B.; Fu, Y. PriSTI: A conditional diffusion framework for spatiotemporal imputation. In Proceedings of the 2023 IEEE 39th International Conference on Data Engineering (ICDE), Anaheim, CA, USA, 11–14 April 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 1927–1939. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Yang, L.; Yang, Y.; Peng, L.; Ge, X. Spatio-temporal graph neural networks for missing data completion in traffic prediction. Int. J. Geogr. Inf. Sci. 2025, 39, 1057–1075. [Google Scholar] [CrossRef] [Scilit]
- Zhan, Z.; Mao, X.; Liu, H.; Yu, S. STGL: Self-supervised spatio-temporal graph learning for traffic forecasting. J. Artif. Intell. Res. 2025, 2, 1–8. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Li, Y.; Zhao, W.; Zhu, H.; Zhang, J.; Wu, X. GSF-LLM: Graph-Enhanced Spatio-Temporal Fusion-Based Large Language Model for Traffic Prediction. Sensors 2025, 25, 6698. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yan, Y.; Liao, Y.; Xu, G.; Yao, R.; Fan, H.; Sun, J.; Wang, X.; Sprinkle, J.; An, Z.; Ma, M.; et al. Large language models for traffic and transportation research: Methodologies, state of the Art, and future opportunities. arXiv 2025, arXiv:2503.21330. [Google Scholar] [CrossRef] [Scilit]
- Scarselli, F.; Gori, M.; Tsoi, A.C.; Hagenbuchner, M.; Monfardini, G. The graph neural network model. IEEE Trans. Neural Netw. 2009, 20, 61–80. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Zhu, J.; Yin, Y.; Ignatius, J.; Wei, X.; Kumar, A. Dynamic travel time prediction with spatiotemporal features: Using a GNN-based deep learning method. Ann. Oper. Res. 2024, 340, 571–591. [Google Scholar] [CrossRef] [Scilit]
- Perez, E.; Strub, F.; de Vries, H.; Dumoulin, V.; Courville, A. FiLM: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018. [Google Scholar] [CrossRef] [Scilit]
- Brockschmidt, M. GNN-FiLM: Graph neural networks with feature-wise linear modulation. In Proceedings of the 37th International Conference on Machine Learning, Virtual, 13–18 July 2020. [Google Scholar] [CrossRef] [Scilit]
- Chen, C.; Liu, Y.; Chen, L.; Zhang, C. Bidirectional spatial-temporal adaptive transformer for urban traffic flow forecasting. IEEE Trans. Neural Netw. Learn. Syst. 2023, 34, 6913–6925. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 6000–6010. [Google Scholar] [CrossRef] [Scilit]
- Pan, Z.; Liang, Y.; Wang, W.; Yu, Y.; Zheng, Y.; Zhang, J. Urban traffic prediction from spatio-temporal data using deep meta learning. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, New York, NY, USA, 4–8 August 2019; ACM: New York, NY, USA, 2019; pp. 1720–1730. [Google Scholar] [CrossRef] [Scilit]
- Tian, Y.; Zhang, K.; Li, J.; Lin, X.; Yang, B. LSTM-based traffic flow prediction with missing data. Neurocomputing 2018, 318, 297–305. [Google Scholar] [CrossRef] [Scilit]
- Bai, S.; Kolter, J.Z.; Koltun, V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv 2018, arXiv:1803.01271. [Google Scholar] [CrossRef] [Scilit]
- Wu, T.; Chen, F.; Wan, Y. Graph Attention LSTM Network: A new model for traffic flow forecasting. In Proceedings of the 2018 5th International Conference on Information Science and Control Engineering (ICISCE), Zhengzhou, China, 20–22 July 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 241–245. [Google Scholar] [CrossRef] [Scilit]
- Bai, L.; Yao, L.; Li, C.; Wang, X.; Wang, C. Adaptive graph convolutional recurrent network for traffic forecasting. Adv. Neural Inf. Process. Syst. 2020, 33, 17804–17815. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Zhang, J.; Cao, W.; Li, J.; Zheng, Y. When will you arrive? Estimating travel time based on deep neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Xu, C.; Zhang, W.; Li, J. GraphTTE: Travel time estimation based on attention-spatiotemporal graphs. IEEE Signal Process. Lett. 2021, 28, 239–243. [Google Scholar] [CrossRef] [Scilit]
- Khaled, A.; Elsir, A.M.T.; Shen, Y. GSTA: Gated spatial-temporal attention approach for travel time prediction. Neural Comput. Appl. 2022, 34, 2307–2322. [Google Scholar] [CrossRef] [Scilit]
- Jiang, W.; Luo, J. Graph neural network for traffic forecasting: A survey. Expert Syst. Appl. 2022, 207, 117921. [Google Scholar] [CrossRef] [Scilit]
- Jin, G.; Liang, Y.; Fang, Y.; Shao, Z.; Huang, J.; Zhang, J.; Zheng, Y. Spatio-temporal graph neural networks for predictive learning in urban computing: A survey. IEEE Trans. Knowl. Data Eng. 2024, 36, 5388–5408. [Google Scholar] [CrossRef] [Scilit]
- Pan, Z.; Zhang, W.; Liang, Y.; Zhang, W.; Yu, Y.; Zhang, J.; Zheng, Y. Spatio-temporal meta learning for urban traffic prediction. IEEE Trans. Knowl. Data Eng. 2022, 34, 1462–1476. [Google Scholar] [CrossRef] [Scilit]
- Ji, J.; Wang, J.; Jiang, Z.; Jiang, J.; Zhang, H. STDEN: Towards physics-guided neural networks for traffic flow prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 22 February–1 March 2022. [Google Scholar] [CrossRef] [Scilit]
- Di, X.; Shi, R.; Mo, Z.; Fu, Y. Physics-Informed Deep Learning for Traffic State Estimation: A Survey and the Outlook. Algorithms 2023, 16, 305. [Google Scholar] [CrossRef] [Scilit]
- Wen, H.; Lin, Y.; Xia, Y.; Wan, H.; Wen, Q.; Zimmermann, R.; Liang, Y. DiffSTG: Probabilistic spatio-temporal graph forecasting with denoising diffusion models. In Proceedings of the 31st ACM International Conference on Advances in Geographic Information Systems, New York, NY, USA, 13–16 November 2023; ACM: New York, NY, USA, 2023; p. 60. [Google Scholar] [CrossRef] [Scilit]
- Liu, D.; Wang, J.; Shang, S.; Han, P. MSDR: Multi-step dependency relation networks for spatial temporal forecasting. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), Washington, DC, USA, 14–18 August 2022; ACM: New York, NY, USA, 2022; pp. 1042–1050. [Google Scholar] [CrossRef] [Scilit]
- Cui, Z.; Henrickson, K.; Ke, R.; Wang, Y. Traffic graph convolutional recurrent neural network: A deep learning framework for network-scale traffic learning and forecasting. IEEE Trans. Intell. Transp. Syst. 2020, 21, 4883–4894. [Google Scholar] [CrossRef] [Scilit]
- Geng, X.; Li, Y.; Wang, L.; Zhang, L.; Yang, Q.; Ye, J.; Liu, Y. Spatiotemporal multi-graph convolution network for ride-hailing demand forecasting. Proc. AAAI Conf. Artif. Intell. 2019, 33, 3656–3663. [Google Scholar] [CrossRef] [Scilit]
- Fang, Z.; Long, Q.; Song, G.; Xie, K. Spatial-temporal graph ODE networks for traffic flow forecasting. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD), Virtual Event, Singapore, 14–18 August 2021; ACM: New York, NY, USA, 2021; pp. 364–373. [Google Scholar] [CrossRef] [Scilit]
- Xu, M.; Dai, W.; Liu, C.; Gao, X.; Lin, W.; Qi, G.-J.; Xiong, H. Spatial-temporal transformer networks for traffic flow forecasting. arXiv 2021, arXiv:2001.02908. [Google Scholar] [CrossRef] [Scilit]
- Guo, K.; Hu, Y.; Sun, Y.; Qian, S.; Gao, J.; Yin, B. Hierarchical graph convolution network for traffic forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual Event, 2–9 February 2021. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Z.; Zhang, X. Dual-view graph convolutional neural networks for urban traffic congestion level prediction. In Proceedings of the 2025 5th International Conference on Neural Networks, Information and Communication Engineering (NNICE), Guangzhou, China, 10–12 January 2025. [Google Scholar] [CrossRef] [Scilit]
- Guo, S.; Lin, Y.; Feng, N.; Song, C.; Wan, H. Attention based spatial-temporal graph convolutional networks for traffic flow forecasting. Proc. AAAI Conf. Artif. Intell. 2019, 33, 922–929. [Google Scholar] [CrossRef] [Scilit]
- Shang, C.; Chen, J.; Bi, J. Discrete graph structure learning for forecasting multiple time series. In Proceedings of the Web Conference (WWW), Ljubljana, Slovenia, 19–23 April 2021. [Google Scholar] [CrossRef] [Scilit]










| Symbol | Description | Dimensions |
|---|---|---|
| G = (V, E) | The road network graph with nodes and edges | – |
| L(G) = (V′, E′) | The Line Graph (Dual Graph) where nodes represent road links | – |
| xstatic, xdyn | Static road attributes and dynamic temporal features | |
| Topological embeddings updated at the k-th GGNN step | ||
| γ, β | Scale and shift parameters generated by FiLM | |
| Hrnn | The global hidden state of the Bi-GRU backbone | |
| Cattn | Context vector aggregated by Dynamic Query Attention | |
| Predicted travel time (log-transformed) | Scalar |
| Parameter Category | Parameter Name | Value |
|---|---|---|
| Model Architecture | ID Embedding Dimension | 32 |
| Hidden Dimension (RNN/GCN) | 64 | |
| Number of GCN Layers | 3 | |
| Attention Heads | 4 | |
| Training Strategy | Optimizer | AdamW |
| Learning Rate (lr) | 1 × 10−3 (ReduceLROnPlateau) | |
| Batch Size | 32 | |
| Loss Function | L1 Loss (Log-Space) | |
| Regularization (Dropout) | 0.3 |
| Variant | MAE (s) | RMSE (s) (Long_RMSE) | MAPE (%) | R2 | Peak_ RMSE | Congested _RMSE | Peak_R2 | Congested_R2 |
|---|---|---|---|---|---|---|---|---|
| LSTM | 488.707 | 869.995 | 49.429 | 0.335 | 787.804 | 541.168 | 0.320 | −0.017 |
| GAT-LSTM | 433.722 | 783.089 | 44.049 | 0.461 | 824.820 | 485.621 | 0.255 | 0.180 |
| DeepTTE | 372.135 | 665.039 | 41.343 | 0.611 | 696.314 | 458.929 | 0.469 | 0.268 |
| TCN | 347.091 | 628.200 | 31.692 | 0.653 | 572.440 | 468.618 | 0.641 | 0.237 |
| GCN-GRU | 337.011 | 623.575 | 32.493 | 0.658 | 627.352 | 464.373 | 0.569 | 0.250 |
| Attn-GRU | 336.483 | 594.758 | 31.350 | 0.689 | 545.636 | 390.303 | 0.674 | 0.470 |
| GraphTTE | 351.558 | 592.563 | 33.413 | 0.691 | 554.076 | 490.899 | 0.664 | 0.162 |
| Graph WaveNet | 293.099 | 556.202 | 28.436 | 0.728 | 480.407 | 324.734 | 0.747 | 0.633 |
| GRU | 293.933 | 548.852 | 31.728 | 0.735 | 550.127 | 370.225 | 0.668 | 0.523 |
| TAGCN-GRU | 280.476 | 518.486 | 30.353 | 0.763 | 506.339 | 399.568 | 0.719 | 0.445 |
| GSTA | 285.916 | 485.312 | 30.902 | 0.793 | 467.265 | 424.048 | 0.761 | 0.375 |
| Adaptive-GGNN | 251.533 | 482.059 | 26.623 | 0.795 | 438.870 | 365.422 | 0.789 | 0.536 |
| DAGRN (Ours) | 225.152 | 415.485 | 24.285 | 0.848 | 401.400 | 377.308 | 0.823 | 0.505 |
| Variant | MAE (s) | RMSE (s) | MAPE (%) | R2 | Long_RMSE (s) | Peak_RMSE (s) | Congested_RMSE (s) |
|---|---|---|---|---|---|---|---|
| DAGRN_no_gcn | 232.441 | 444.166 | 26.384 | 0.826 | 444.166 | 447.681 | 345.409 |
| DAGRN_no_attn | 287.816 | 521.495 | 32.783 | 0.761 | 521.495 | 493.713 | 429.464 |
| DAGRN_no_Stratification | 239.268 | 429.366 | 24.468 | 0.838 | 429.366 | 426.544 | 431.177 |
| DAGRN_no_film | 230.421 | 447.665 | 25.428 | 0.823 | 447.665 | 480.111 | 352.867 |
| DAGRN_no_gate | 239.996 | 454.289 | 27.593 | 0.818 | 454.289 | 473.743 | 357.867 |
| DAGRN_Full | 225.152 | 415.485 | 24.285 | 0.848 | 415.485 | 401.400 | 377.308 |
| Parameter Changed | Hidden Dim | Dropout Rate | Learning Rate | GCN Iter | Attention Heads | Batch Size | RMSE (s) | MAE (s) |
|---|---|---|---|---|---|---|---|---|
| hidden_dim = 32 | 32 | – | – | – | – | – | 476.20 | 246.26 |
| hidden_dim = 64 | 64 | – | – | – | – | – | 430.75 | 220.75 |
| hidden_dim = 128 | 128 | – | – | – | – | – | 402.59 | 221.48 |
| dropout_rate = 0.2 | – | 0.2 | – | – | – | – | 401.95 | 216.12 |
| dropout_rate = 0.5 | – | 0.5 | – | – | – | – | 433.74 | 238.63 |
| dropout_rate = 0.7 | – | 0.7 | – | – | – | – | 534.12 | 291.17 |
| lr = 0.0005 | – | – | 0.0005 | – | – | – | 466.16 | 232.73 |
| lr = 0.001 | – | – | 0.001 | – | – | – | 399.92 | 217.10 |
| lr = 0.005 | – | – | 0.005 | – | – | – | 457.78 | 263.71 |
| gcn_iter = 1 | – | – | – | 1 | – | – | 467.34 | 241.53 |
| gcn_iter = 2 | – | – | – | 2 | – | – | 446.32 | 229.24 |
| gcn_iter = 3 | – | – | – | 3 | – | – | 444.81 | 232.15 |
| attention_heads = 2 | – | – | – | – | 2 | – | 425.65 | 220.82 |
| attention_heads = 4 | – | – | – | – | 4 | – | 419.26 | 225.66 |
| attention_heads = 8 | – | – | – | – | 8 | – | 491.83 | 252.03 |
| batch_size = 32 | – | – | – | – | – | 32 | 399.89 | 214.22 |
| batch_size = 64 | – | – | – | – | – | 64 | 442.78 | 228.58 |
| batch_size = 128 | – | – | – | – | – | 128 | 479.08 | 245.85 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Published by MDPI on behalf of the International Society for Photogrammetry and Remote Sensing. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wu, J.; Zhang, Y.; Bai, Y.; Xia, J.; He, Y. A Dual-Attentional Gated Residual Framework for Robust Travel Time Prediction. ISPRS Int. J. Geo-Inf. 2026, 15, 120. https://doi.org/10.3390/ijgi15030120
Wu J, Zhang Y, Bai Y, Xia J, He Y. A Dual-Attentional Gated Residual Framework for Robust Travel Time Prediction. ISPRS International Journal of Geo-Information. 2026; 15(3):120. https://doi.org/10.3390/ijgi15030120
Chicago/Turabian StyleWu, Jiajun, Yongchuan Zhang, Yiduo Bai, Jun Xia, and Yong He. 2026. "A Dual-Attentional Gated Residual Framework for Robust Travel Time Prediction" ISPRS International Journal of Geo-Information 15, no. 3: 120. https://doi.org/10.3390/ijgi15030120
APA StyleWu, J., Zhang, Y., Bai, Y., Xia, J., & He, Y. (2026). A Dual-Attentional Gated Residual Framework for Robust Travel Time Prediction. ISPRS International Journal of Geo-Information, 15(3), 120. https://doi.org/10.3390/ijgi15030120

