PHR-Net: Proposal-Level Historical Retrieval for Non-Stationary Temporal Consistency in Trajectory Prediction
Abstract
1. Introduction
2. Related Work
2.1. Classical Trajectory Modeling and Interaction Learning
2.2. Vectorized Scene Representation
2.3. Multimodal Outputs
2.4. Query-Based Modeling
2.5. Non-Stationary Temporal Consistency
3. Method
3.1. Problem Formulation
3.2. Overall Architecture
3.3. Scene Attention
3.4. Agent Attention
3.5. Mode Attention
3.6. Stage-2: Intention Branch
3.7. Stage-2: Trajectory Branch
3.8. Training Objective
4. Experimentation and Evaluation
4.1. Experimental Setup
4.1.1. Dataset
4.1.2. Evaluation Metrics
4.1.3. Implementation Details
4.2. Experimental Results
4.2.1. Quantitative Results
- LaneGCN [9]: A representative lane-graph modeling method. This method constructs a lane graph based on vectorized high-definition maps and models information propagation between lanes, between lanes and vehicles, and between vehicles through graph convolution.
- DenseTNT [30]: An anchor-free trajectory prediction method. This method estimates the probabilities of all dense candidate goal points on the road, then selects high-confidence goals through non-maximum suppression, and finally completes a full trajectory for each selected goal.
- Scene Transformer [23]: A scene-level Transformer prediction method. This method fuses map elements, historical trajectories, agent interactions, and temporal context information through attention mechanisms to achieve joint multi-agent trajectory prediction.
- HiVT [10]: A hierarchical vectorized Transformer method. This method divides the prediction process into two levels: local context extraction and global interaction modeling. It also introduces translation-invariant representation and rotation-invariant spatial learning modules to improve modeling efficiency and geometric robustness.
- Wayformer [25]: A trajectory prediction method based on a unified attention structure. This method adopts a homogeneous Transformer backbone to fuse historical trajectories, map semantics, and traffic participant interaction information, reducing handcrafted module design and improving multi-source scene information modeling capability.
- DCMS [18]: A training-enhancement method for continuous prediction stability. This method suppresses trajectory deviations under adjacent predictions or perturbed inputs through temporal consistency and spatial consistency constraints and uses multi-pseudo-target supervision to enhance multimodal learning capability.
- HPNet [19]: A modeling method that treats historical embedding representations as inputs for future timesteps. It introduces a historical prediction attention mechanism, using the agent interaction results from previous timesteps as reference information for the current prediction, thereby improving cross-timestep prediction consistency and stability.
- QCNet [16]: A trajectory prediction method based on the query-centric idea. This method reduces repeated computation in sliding-window online prediction through query-centric scene encoding and adopts a two-stage decoding structure that combines trajectory proposal generation and refinement to improve multimodal prediction accuracy and online inference efficiency.
4.2.2. Qualitative Results
4.3. Ablation Studies
4.3.1. Ablation Settings
4.3.2. Sensitivity Analysis
4.3.3. Operational Efficiency Analysis
4.3.4. Consistency Verification on the Overlapping Portions of Adjacent Predictions
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| MLP | Multi-Layer Perceptron |
| CNN | Convolutional Neural Network |
| GNN | Graph Neural Network |
| GRU | Gated Recurrent Unit |
| HD | High-Definition |
| GT | Ground Truth |
| AdamW | Adaptive Moment Estimation with decoupled Weight Decay |
| minADE | Minimum Average Displacement Error |
| minFDE | Minimum Final Displacement Error |
| MR | Miss Rate |
| summed ADE | Summed Average Displacement Error |
| FPS | Frames Per Second |
| Top-1 | Evaluation using only the highest-confidence predicted trajectory |
| Top-6 | Evaluation using the best-matched trajectory among the top six predicted trajectories |
Appendix A. Detailed Flowchart
| Compact Notation in Flowchart | Full Notation or Meaning in the Main Text |
|---|---|
| Historical trajectory of target agent i | |
| Neighboring-agent set of target agent i | |
| Local map elements of target agent i | |
| Scene-attention feature | |
| Agent-attention feature | |
| Mode-attention feature | |
| Stage-1 coarse proposal | |
| B | Historical proposal memory of target agent i |
| g | High-level behavior representation |
| Prototype matching weight | |
| Retrieved intention-history feature | |
| Second-stage mode feature | |
| s | Proposal sequence encoding |
| Retrieved trajectory-history context | |
| Trajectory-history feature | |
| o | Fused feature |
| Residual correction | |
| Final refined trajectory | |
| Mode probability | |
| Total training objective |

References
- Kalman, R.E. A New Approach to Linear Filtering and Prediction Problems. J. Basic Eng. 1960, 82, 35–45. [Google Scholar] [CrossRef] [Scilit]
- Helbing, D.; Molnár, P. Social force model for pedestrian dynamics. Phys. Rev. E 1995, 51, 4282–4286. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alahi, A.; Goel, K.; Ramanathan, V.; Robicquet, A.; Fei-Fei, L.; Savarese, S. Social LSTM: Human Trajectory Prediction in Crowded Spaces. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 961–971. [Google Scholar] [CrossRef] [Scilit]
- Lee, N.; Choi, W.; Vernaza, P.; Choy, C.B.; Torr, P.H.S.; Chandraker, M. DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2165–2174. [Google Scholar] [CrossRef] [Scilit]
- Gupta, A.; Johnson, J.; Fei-Fei, L.; Savarese, S.; Alahi, A. Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 2255–2264. [Google Scholar] [CrossRef] [Scilit]
- Ivanovic, B.; Pavone, M. The Trajectron: Probabilistic Multi-Agent Trajectory Modeling with Dynamic Spatiotemporal Graphs. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, South Korea, 27 October–2 November 2019; pp. 2375–2384. [Google Scholar] [CrossRef] [Scilit]
- Salzmann, T.; Ivanovic, B.; Chakravarty, P.; Pavone, M. Trajectron++: Dynamically-Feasible Trajectory Forecasting with Heterogeneous Data. In Computer Vision—ECCV 2020; Vedaldi, A., Bischof, H., Brox, T., Frahm, J.M., Eds.; Springer International Publishing: Cham, Switzerland, 2020; pp. 683–700. [Google Scholar] [CrossRef] [Scilit]
- Gao, J.; Sun, C.; Zhao, H.; Shen, Y.; Anguelov, D.; Li, C.; Schmid, C. VectorNet: Encoding HD Maps and Agent Dynamics From Vectorized Representation. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 14–19 June 2020; pp. 11522–11530. [Google Scholar] [CrossRef] [Scilit]
- Liang, M.; Yang, B.; Hu, R.; Chen, Y.; Liao, R.; Feng, S.; Urtasun, R. Learning Lane Graph Representations for Motion Forecasting. In Computer Vision—ECCV 2020; Vedaldi, A., Bischof, H., Brox, T., Frahm, J.M., Eds.; Springer International Publishing: Cham, Switzerland, 2020; pp. 541–556. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.; Ye, L.; Wang, J.; Wu, K.; Lu, K. HiVT: Hierarchical Vector Transformer for Multi-Agent Motion Prediction. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 8813–8823. [Google Scholar] [CrossRef] [Scilit]
- Jia, X.; Wu, P.; Chen, L.; Liu, Y.; Li, H.; Yan, J. HDGT: Heterogeneous Driving Graph Transformer for Multi-Agent Trajectory Prediction via Scene Encoding. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13860–13875. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chai, Y.; Sapp, B.; Bansal, M.; Anguelov, D. MultiPath: Multiple Probabilistic Anchor Trajectory Hypotheses for Behavior Prediction. In Proceedings of the Conference on Robot Learning; Kaelbling, L.P., Kragic, D., Sugiura, K., Eds.; (Proceedings of Machine Learning Research); PMLR: Online, 2020; Volume 100, pp. 86–99. Available online: https://proceedings.mlr.press/v100/chai20a.html (accessed on 10 May 2026).
- Phan-Minh, T.; Grigore, E.C.; Boulton, F.A.; Beijbom, O.; Wolff, E.M. CoverNet: Multimodal Behavior Prediction Using Trajectory Sets. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 14–19 June 2020; pp. 14062–14071. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Gao, J.; Lan, T.; Sun, C.; Sapp, B.; Varadarajan, B.; Shen, Y.; Shen, Y.; Chai, Y.; Schmid, C.; et al. TNT: Target-driven Trajectory Prediction. In Proceedings of the 2020 Conference on Robot Learning; Kober, J., Ramos, F., Tomlin, C., Eds.; (Proceedings of Machine Learning Research); PMLR: Online, 2021; Volume 155, pp. 895–904. Available online: https://proceedings.mlr.press/v155/zhao21b.html (accessed on 10 May 2026).
- Shi, S.; Jiang, L.; Dai, D.; Schiele, B. Motion Transformer with Global Intention Localization and Local Movement Refinement. In Proceedings of the Advances in Neural Information Processing Systems; Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2022; pp. 6531–6543. Available online: https://proceedings.neurips.cc/paper_files/paper/2022/file/2ab47c960bfee4f86dfc362f26ad066a-Paper-Conference.pdf (accessed on 10 May 2026).
- Zhou, Z.; Wang, J.; Li, Y.H.; Huang, Y.K. Query-Centric Trajectory Prediction. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 17863–17873. [Google Scholar] [CrossRef] [Scilit]
- Jiang, C.M.; Cornman, A.; Park, C.; Sapp, B.; Zhou, Y.; Anguelov, D. MotionDiffuser: Controllable Multi-Agent Motion Prediction Using Diffusion. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 9644–9653. [Google Scholar] [CrossRef] [Scilit]
- Ye, M.; Xu, J.; Xu, X.; Cao, T.; Chen, Q. DCMS: Motion Forecasting with Dual Consistency and Multi-Pseudo-Target Supervision. arXiv 2022, arXiv:2204.05859. [Google Scholar]
- Tang, X.; Kan, M.; Shan, S.; Ji, Z.; Bai, J.; Chen, X. HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 15261–15270. [Google Scholar] [CrossRef] [Scilit]
- Zang, Z.; Zhang, X.; Gong, X.; Song, J.; Yu, R.; Gong, J. A spatio-temporal trajectory planning framework for AGVs based on motion primitive and dynamic programming in off-road environments. Adv. Eng. Inform. 2026, 69, 103934. [Google Scholar] [CrossRef] [Scilit]
- Casas, S.; Luo, W.; Urtasun, R. IntentNet: Learning to Predict Intention from Raw Sensor Data. In Proceedings of the 2nd Conference on Robot Learning; Billard, A., Dragan, A., Peters, J., Morimoto, J., Eds.; PMLR: Online, 2018; Volume 87, pp. 947–956. Available online: https://proceedings.mlr.press/v87/casas18a.html (accessed on 10 May 2026).
- Hong, J.; Sapp, B.; Philbin, J. Rules of the Road: Predicting Driving Behavior with a Convolutional Model of Semantic Interactions. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 8446–8454. [Google Scholar] [CrossRef] [Scilit]
- Ngiam, J.; Vasudevan, V.; Caine, B.; Zhang, Z.; Chiang, H.T.L.; Ling, J.; Roelofs, R.; Bewley, A.; Liu, C.; Venugopal, A.; et al. Scene Transformer: A unified architecture for predicting future trajectories of multiple agents. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022; Available online: https://openreview.net/forum?id=Wm3EA5OlHs (accessed on 10 May 2026).
- Yuan, Y.; Weng, X.; Ou, Y.; Kitani, K. AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 9793–9803. [Google Scholar] [CrossRef] [Scilit]
- Nayakanti, N.; Al-Rfou, R.; Zhou, A.; Goel, K.; Refaat, K.S.; Sapp, B. Wayformer: Motion forecasting via simple and efficient attention networks. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 2980–2987. [Google Scholar] [CrossRef] [Scilit]
- Sun, Q.; Huang, X.; Gu, J.; Williams, B.C.; Zhao, H. M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 6533–6542. [Google Scholar] [CrossRef] [Scilit]
- Marchetti, F.; Becattini, F.; Seidenari, L.; Del Bimbo, A. MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 14–19 June 2020; pp. 7141–7150. [Google Scholar] [CrossRef] [Scilit]
- Rhinehart, N.; Kitani, K.M.; Vernaza, P. r2p2: A ReparameteRized Pushforward Policy for Diverse, Precise Generative Path Forecasting. In Computer Vision – ECCV 2018; Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y., Eds.; Springer International Publishing: Cham, Switzerland, 2018; pp. 794–811. [Google Scholar] [CrossRef] [Scilit]
- Chang, M.F.; Lambert, J.; Sangkloy, P.; Singh, J.; Bak, S.; Hartnett, A.; Wang, D.; Carr, P.; Lucey, S.; Ramanan, D.; et al. Argoverse: 3D Tracking and Forecasting with Rich Maps. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 8740–8749. [Google Scholar] [CrossRef] [Scilit]
- Gu, J.; Sun, C.; Zhao, H. DenseTNT: End-to-end Trajectory Prediction from Dense Goal Sets. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 15283–15292. [Google Scholar] [CrossRef] [Scilit]






| Component | Specification |
|---|---|
| GPU | GeForce RTX 4090 |
| CPU | AMD EPYC 7K62 48 |
| Memory | 64 GB DDR5 |
| Storage | 2 TB SSD |
| Operating system | Ubuntu 20.04 |
| CUDA version | 12.4 |
| Deep learning framework | PyTorch 2.1.0 |
| Method | minADE1 ↓ | minFDE1 ↓ | MR1 ↓ |
|---|---|---|---|
| LaneGCN [9] | 0.8703 | 1.3622 | 0.1620 |
| DenseTNT [30] | 0.8817 | 1.2815 | 0.1258 |
| Scene Transformer [23] | 0.8026 | 1.2321 | 0.1255 |
| HiVT [10] | 0.7735 | 1.1693 | 0.1267 |
| HPNet [19] | 0.7612 | 1.0986 | 0.1067 |
| PHR-Net | 0.7635 | 1.0834 | 0.1046 |
| Method | minADE6 ↓ | minFDE6 ↓ | MR6 ↓ |
|---|---|---|---|
| Wayformer [25] | 0.7676 | 1.1616 | 0.1186 |
| DCMS [18] | 0.7659 | 1.1350 | 0.1094 |
| QCNet [16] | 0.7340 | 1.0666 | 0.1056 |
| PHR-Net | 0.7432 | 1.0804 | 0.1027 |
| Model Configuration | Intent Branch | Trajectory Branch | Static Prototype Library | minADE ↓ | minFDE ↓ | MR ↓ |
|---|---|---|---|---|---|---|
| Base | 0.778 | 1.185 | 0.128 | |||
| Base + Intent | ✓ | ✓ | 0.733 | 1.024 | 0.084 | |
| Base + Traj | ✓ | 0.675 | 0.976 | 0.106 | ||
| Intent + Traj w/o Proto | ✓ | ✓ | 0.670 | 0.937 | 0.094 | |
| PHR-Net | ✓ | ✓ | ✓ | 0.641 | 0.858 | 0.067 |
| Method | Params (M) ↓ | Latency (ms) ↓ | FPS ↑ | Peak Memory (GB) ↓ |
|---|---|---|---|---|
| HPNet | 4.070 | 41.841 | 23.900 | 1.770 |
| PHR-Net | 3.601 | 48.372 | 20.673 | 1.932 |
| Method | Summed ADE ↓ |
|---|---|
| Baseline w/o historical reference | 3.00 |
| HPNet [19] | 2.25 |
| PHR-Net | 2.08 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, B.; Xu, M. PHR-Net: Proposal-Level Historical Retrieval for Non-Stationary Temporal Consistency in Trajectory Prediction. Vehicles 2026, 8, 109. https://doi.org/10.3390/vehicles8050109
Zhang B, Xu M. PHR-Net: Proposal-Level Historical Retrieval for Non-Stationary Temporal Consistency in Trajectory Prediction. Vehicles. 2026; 8(5):109. https://doi.org/10.3390/vehicles8050109
Chicago/Turabian StyleZhang, Bo, and Ming Xu. 2026. "PHR-Net: Proposal-Level Historical Retrieval for Non-Stationary Temporal Consistency in Trajectory Prediction" Vehicles 8, no. 5: 109. https://doi.org/10.3390/vehicles8050109
APA StyleZhang, B., & Xu, M. (2026). PHR-Net: Proposal-Level Historical Retrieval for Non-Stationary Temporal Consistency in Trajectory Prediction. Vehicles, 8(5), 109. https://doi.org/10.3390/vehicles8050109
