Hybrid Attribution-Based Interpretable Deep Reinforcement Learning for Autonomous Driving Behavior Decision-Making
Abstract
1. Introduction
- We extend gradient-based sensitivity analysis to the KAN-based value function and integrate it with LRP-based contribution allocation, thereby proposing a hybrid attribution mechanism. The proposed method provides computationally efficient, real-time, comprehensive, and semantically aligned explanations for DRL decision-making.
- We introduce KAN into the Dueling DQN framework to enhance structural interpretability and improve the consistency of attribution propagation.
- We validate the proposed HA-IDRL framework on a highway lane-changing task, demonstrating that it achieves competitive driving performance while significantly improving interpretability quality.
2. Interpretable Deep Reinforcement Learning Methodology
2.1. HA-IDRL Framework Overview
2.2. Kolmogorov–Arnold Networks for Value Function Modeling
2.3. Gradient-Based Attribution with KAN
2.4. Hybrid Gradient–LRP Attribution Method
| Algorithm 1 Training and Explanation Procedure of HA-IDRL |
| Require: Environment Env, replay buffer M, maximum episodes T, mini-batch size N, discount factor γ Ensure: Trained KAN-based Dueling DQN model Q(s, a) and hybrid attribution result HGL 1: Initialize environment Env, KAN-based Dueling DQN evaluation network Qθ(s, a), target network Qϕ(s, a), and replay buffer M 2: for episode = 1 to T do 3: Reset Env and observe initial state s0 4: while episode not terminated do 5: Select action at using ε-greedy policy based on Qθ(st, a) 6: Execute action at in Env and observe reward rt, next state st+1, terminal flag dt 7: Store transition (st, at, rt, st+1, dt) in M 8: if M is ready for training then 9: Sample a mini-batch {(si, ai, ri, si+1, di)} from M 10: Compute target value: yi = ri + γ(1 − di) maxa′Qϕ(si+1, a′) and TD loss: L = (1/N) Σi (Qθ(si, ai) − yi)2 11: Update θ by minimizing L 12: end if 13: st ← st+1 14: Periodically update target network ϕ ← θ 15: end while 16: end for 17: Given a trained KAN-based Dueling DQN model and an input state s 18: Compute gradient attribution g for the explained action 19: Compute LRP attribution r for the same decision target 20: Fuse g and r using the HGL rule to obtain I 21: Output I as the final feature-level explanation |
3. Experimental Evaluation
3.1. Simulation Environment
3.2. State and Action Spaces
3.3. Reward Design
3.3.1. Speed Reward
3.3.2. Safety Penalty
3.3.3. Energy Consumption Penalty
3.3.4. Comfort Penalty
3.4. Baselines and Training Configuration
- PPO: PPO is an on-policy actor–critic algorithm that stabilizes training via a clipped objective.
- SAC: SAC is an off-policy maximum entropy method that balances exploration and exploitation.
- DDQN: DDQN mitigates Q-value overestimation through decoupled action selection and evaluation.
- Dueling DQN: Dueling DQN improves learning efficiency by decomposing Q-values into state value and action advantage components.
3.5. Decision-Making Performance Evaluation
4. Interpretability Analysis
4.1. Local Explanation Analysis
4.2. Global Explanation Analysis
4.3. Summary of Experimental Results
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| DRL | Deep Reinforcement Deep |
| DQN | Deep Q-Network |
| DDQN | Double Deep Q-Network |
| Dueling DQN | Dueling Deep Q-Network |
| KAN | Kolmogorov–Arnold Networks |
| MLP | multilayer perceptron |
| LRP | Layer-wise Relevance Propagation |
| HGL | Hybrid Gradient–LRP Attribution |
| HA-IDRL | Hybrid Attribution-based Interpretable Deep Reinforcement Learning |
References
- Paden, B.; Čáp, M.; Yong, S.Z.; Yershov, D.; Frazzoli, E. A survey of motion planning and control techniques for self-driving urban vehicles. IEEE Trans. Intell. Veh. 2016, 1, 33–55. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Xi, Z.; Xu, Y.; Xiong, Z.; Ding, X.; Wang, H. Balancing safety and efficiency for autonomous vehicles at urban uncontrolled crosswalk: Challenges and countermeasures. Accid. Anal. Prev. 2025, 220, 108111. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kurt, A.; Özgüner, Ü. Hierarchical finite state machines for autonomous mobile systems. Control. Eng. Pract. 2013, 21, 184–194. [Google Scholar] [CrossRef] [Scilit]
- Dolgov, D.; Thrun, S.; Montemerlo, M.; Diebel, J. Path planning for autonomous vehicles in unknown semi-structured environments. Int. J. Robot. Res. 2010, 29, 485–501. [Google Scholar] [CrossRef] [Scilit]
- Kesting, A.; Treiber, M.; Helbing, D. General lane-changing model MOBIL for car-following models. Transp. Res. Rec. 2007, 1999, 86–94. [Google Scholar] [CrossRef] [Scilit]
- Bojarski, M.; Del Testa, D.; Dworakowski, D.; Firner, B.; Flepp, B.; Goyal, P.; Jackel, L.D.; Monfort, M.; Muller, U.; Zhang, J.; et al. End to end learning for self-driving cars. arXiv 2016, arXiv:1604.07316. [Google Scholar] [CrossRef] [Scilit]
- Codevilla, F.; Müller, M.; López, A.; Koltun, V.; Dosovitskiy, A. End-to-end driving via conditional imitation learning. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21–25 May 2018; IEEE Press: New York, NY, USA, 2018; pp. 4693–4700. [Google Scholar]
- Codevilla, F.; Lopez, A.M.; Koltun, V.; Dosovitskiy, A. On offline evaluation of vision-based driving models. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Berlin/Heidelberg, Germany, 2018; pp. 236–251. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016; AAAI Press: Washington, DC, USA, 2016; Volume 30. [Google Scholar]
- Wang, Z.; Schaul, T.; Hessel, M.; Hasselt, H.V.; Lanctot, M.; Freitas, N.D.; Claims, A.I. Dueling network architectures for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning PMLR, New York, NY, USA, 20–22 June 2016; pp. 1995–2003. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden, 10–15 July 2018; pp. 1861–1870. [Google Scholar]
- Shalev-Shwartz, S.; Shammah, S.; Shashua, A. Safe, multi-agent, reinforcement learning for autonomous driving. arXiv 2016, arXiv:1610.03295. [Google Scholar] [CrossRef] [Scilit]
- Salmani Pour Avval, S.; Eskue, N.D.; Groves, R.M.; Yaghoubi, V. Systematic review on neural architecture search. Artif. Intell. Rev. 2025, 58, 73. [Google Scholar] [CrossRef] [Scilit]
- Su, P.; Xiang, C.; Chen, D. Adopting graph neural networks to understand and reason about dynamic driving scenarios. IEEE Open J. Intell. Transp. Syst. 2025, 6, 579–589. [Google Scholar] [CrossRef] [Scilit]
- Miao, Q.; Jia, L.; Xie, K.; Fu, K.; Yang, Z. A Comprehensive Survey and Taxonomy of Mamba: Applications, Challenges, and Future Directions. Inf. Fusion 2025, 130, 104094. [Google Scholar] [CrossRef] [Scilit]
- Du, S.; Zhu, Z.; Wang, X.; Han, H.; Qiao, J. Real-time local path planning strategy based on deep distributional reinforcement learning. Neurocomputing 2024, 599, 128085. [Google Scholar] [CrossRef] [Scilit]
- Fujimoto, S.; Chang, W.D.; Smith, E.; Gu, S.S.; Precup, D.; Meger, D. For sale: State-action representation learning for deep reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 61573–61624. [Google Scholar]
- Ali, Y.; Hussain, F.; Bliemer, M.C.J.; Zheng, Z.; Haque, M. Predicting and explaining lane-changing behaviour using machine learning: A comparative study. Transp. Res. Part C Emerg. Technol. 2022, 145, 103931. [Google Scholar] [CrossRef] [Scilit]
- Arrieta, A.B.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; Garcia, S.; Gil-Lopez, S.; Molina, D.; Benjamins, R.; et al. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef] [Scilit]
- Hassija, V.; Chamola, V.; Mahapatra, A.; Singal, A.; Goel, D.; Huang, K.; Scardapane, S.; Spinelli, I.; Mahmud, M.; Hussain, A. Interpreting black-box models: A review on explainable artificial intelligence. Cogn. Comput. 2024, 16, 45–74. [Google Scholar] [CrossRef] [Scilit]
- Harinarayan, R.R.A.; Shalinie, S.M. XFDDC: EXplainable Fault Detection Diagnosis and Correction framework for chemical process systems. Process Saf. Environ. Prot. 2022, 165, 463–474. [Google Scholar] [CrossRef] [Scilit]
- Lisboa, P.J.G.; Saralajew, S.; Vellido, A.; Fernández-Domenech, R.; Villmann, T. The coming of age of interpretable and explainable machine learning models. Neurocomputing 2023, 535, 25–39. [Google Scholar] [CrossRef] [Scilit]
- Mathew, D.E.; Ebem, D.U.; Ikegwu, A.C.; Ukeoma, P.E.; Dibiaezue, N.F. Recent emerging techniques in explainable artificial intelligence to enhance the interpretable and understanding of AI models for human. Neural Process. Lett. 2025, 57, 16. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
- Omeiza, D.; Webb, H.; Jirotka, M.; Kunze, L. Explanations in autonomous driving: A survey. IEEE Trans. Intell. Transp. Syst. 2021, 23, 10142–10162. [Google Scholar] [CrossRef] [Scilit]
- Doshi-Velez, F.; Kim, B. Towards A Rigorous Science of Interpretable Machine Learning. arXiv 2017, arXiv:1702.08608. [Google Scholar] [CrossRef] [Scilit]
- Bastani, O.; Kim, C.; Bastani, H. Interpreting blackbox models via model extraction. arXiv 2017, arXiv:1705.08504. [Google Scholar]
- Bastani, O.; Inala, J.P.; Solar-Lezama, A. Interpretable, verifiable, and robust reinforcement learning via program synthesis. In xxAI—Beyond Explainable AI: International Workshop, Held in Conjunction with ICML 2020, Vienna, Austria, 18 July 2020, Revised and Extended Papers; Springer International Publishing: Cham, Switzerland, 2020; pp. 207–228. [Google Scholar]
- Kenny, E.M.; Tucker, M.; Shah, J. Towards interpretable deep reinforcement learning with human-friendly prototypes. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should I trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; ACM: New York, NY, USA, 2016; pp. 1135–1144. [Google Scholar]
- Scott M. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17), Long Beach, CA, USA, 4–9 December 2017; Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 4768–4777. [Google Scholar]
- Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic attribution for deep networks. In Proceedings of the International Conference on Machine Learning, PMLR, Sydney, Australia, 6–11 August 2017; pp. 3319–3328. [Google Scholar]
- Bach, S.; Binder, A.; Montavon, G.; Klauschen, F.; Müller, K.-R.; Samek, W. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE 2015, 10, e0130140. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, J.; Zhou, R.; Li, M.; Li, H.; Liu, Y.; Song, X. From black-box to white-box: Interpretable deep reinforcement learning with Kolmogorov-Arnold networks for autonomous driving. Transp. Res. Part C Emerg. Technol. 2026, 182, 105386. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Soljačić, M.; Hou, T.Y.; Tegmark, M. Kan: Kolmogorov-Arnold networks. arXiv 2024, arXiv:2404.19756. [Google Scholar]
- Leurent, E. An Environment for Autonomous Driving Decision-Making[EB/OL]. GitHub Repository. 2018. Available online: https://github.com/eleurent/highway-env (accessed on 3 February 2026).
- Wang, G.; Hu, J.; Li, Z.; Li, L. Harmonious lane changing via deep reinforcement learning. IEEE Trans. Intell. Transp. Syst. 2021, 23, 4642–4650. [Google Scholar] [CrossRef] [Scilit]









| Model | MLP | KAN |
|---|---|---|
| Recurrence Formula | ||
| Gradient Attribution | ||
| LRP | ||
| Multi Layer Framework | ![]() | ![]() |
| Parameter | Value |
|---|---|
| Discount Factor | 0.8 |
| Learning Rate | 0.0003 |
| Batch Size | 256 |
| MLP Hidden Units | 256 |
| Soft Update Rate | 0.005 |
| Training Update Capacity | 0.005 |
| Replay Buffer Capacity | 106 |
| Grid size (number of knots) | 5 |
| Spline Order | 3 |
| KAN Layer Width | 11 |
| Parameter | Return | Success Rate | Avg Speed | Safety | Comfort | Fuel |
|---|---|---|---|---|---|---|
| PPO | 20.88 ± 0.09 | 0.95 ± 0.01 | 20.08 ± 0.03 | −0.05 ± 0.03 | −0.06 ± 0.004 | −0.07 ± 0.007 |
| SAC | 19.21 ± 0.02 | 0.94 ± 0.02 | 21.13 ± 0.08 | −0.71 ± 0.09 | −2.44 ± 0.07 | −0.75 ± 0.04 |
| DDQN | 21.49 ± 1.08 | 0.993 ± 0.006 | 20.62 ± 0.08 | −0.04 ± 0.04 | −0.40 ± 0.10 | −0.13 ± 0.07 |
| Dueling DQN | 21.51 ± 0.12 | 0.991 ± 0.005 | 20.42 ± 0.03 | −0.03 ± 0.01 | −0.30 ± 0.02 | −0.12 ± 0.01 |
| HA-IDRL (Ours) | 21.98 ± 0.20 | 0.993 ± 0.005 | 20.44 ± 0.015 | −0.015 ± 0.01 | −0.28 ± 0.02 | −0.17 ± 0.018 |
| Parameter | Return | Success Rate | Avg Speed | Safety | Comfort | Fuel |
|---|---|---|---|---|---|---|
| PPO | 29.80 ± 0.66 | 0.990 ± 0.01 | 23.82 ± 0.07 | −0.03 ± 0.03 | −0.13 ± 0.01 | −0.10 ± 0.04 |
| SAC | 37.99 ± 1.02 | 0.95 ± 0.03 | 28.26 ± 0.06 | −0.07 ± 0.04 | −2.58 ± 0.07 | −0.69 ± 0.04 |
| DDQN | 37.22 ± 0.32 | 0.993 ± 0.01 | 26.49 ± 0.15 | −0.002 ± 0.004 | −0.40 ± 0.01 | −0.15 ± 0.01 |
| Dueling DQN | 40.42 ± 0.27 | 0.98 ± 0.02 | 28.08 ± 0.15 | −0.02 ± 0.01 | −0.55 ± 0.02 | −0.17 ± 0.02 |
| HA-IDRL (Ours) | 40.93 ± 0.59 | 0.995 ± 0.01 | 27.79 ± 0.09 | −0.001 ± 0.002 | −0.59 ± 0.03 | −0.18 ± 0.01 |
| Parameter | Return | Success Rate | Avg Speed | Safety | Comfort | Fuel |
|---|---|---|---|---|---|---|
| PPO | 26.14 ± 0.20 | 0.98 ± 0.01 | 22.16 ± 0.05 | −0.02 ± 0.01 | −0.08 ± 0.003 | −0.07 ± 0.01 |
| SAC | 28.50 ± 0.20 | 0.95 ± 0.03 | 24.73 ± 0.06 | −0.22 ± 0.02 | −2.55 ± 0.02 | −0.92 ± 0.02 |
| DDQN | 28.36 ± 0.49 | 0.98 ± 0.02 | 23.04 ± 0.21 | −0.02 ± 0.03 | −0.03 ± 0.06 | −0.14 ± 0.02 |
| Dueling DQN | 29.15 ± 0.08 | 0.990 ± 0.01 | 23.77 ± 0.07 | −0.01 ± 0.01 | −0.56 ± 0.02 | −0.23 ± 0.01 |
| HA-IDRL (Ours) | 29.69 ± 0.30 | 0.992 ± 0.01 | 23.55 ± 0.06 | −0.02 ± 0.01 | −0.49 ± 0.02 | −0.22 ± 0.01 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, Y.; Huang, J.; Li, M.; Ye, Q.; Song, X. Hybrid Attribution-Based Interpretable Deep Reinforcement Learning for Autonomous Driving Behavior Decision-Making. Appl. Sci. 2026, 16, 3096. https://doi.org/10.3390/app16063096
Liu Y, Huang J, Li M, Ye Q, Song X. Hybrid Attribution-Based Interpretable Deep Reinforcement Learning for Autonomous Driving Behavior Decision-Making. Applied Sciences. 2026; 16(6):3096. https://doi.org/10.3390/app16063096
Chicago/Turabian StyleLiu, Yaxuan, Jiakun Huang, Mingjun Li, Qing Ye, and Xiaolin Song. 2026. "Hybrid Attribution-Based Interpretable Deep Reinforcement Learning for Autonomous Driving Behavior Decision-Making" Applied Sciences 16, no. 6: 3096. https://doi.org/10.3390/app16063096
APA StyleLiu, Y., Huang, J., Li, M., Ye, Q., & Song, X. (2026). Hybrid Attribution-Based Interpretable Deep Reinforcement Learning for Autonomous Driving Behavior Decision-Making. Applied Sciences, 16(6), 3096. https://doi.org/10.3390/app16063096



