Attribution-Guided Active Exploration in Deep Reinforcement Learning for Autonomous Driving Decision-Making
Abstract
1. Introduction
- We propose an intrinsic exploration signal based on attribution sensitivity. By applying small perturbations to the state and measuring the Euclidean variation in action-wise attributions, the proposed method identifies actions whose decision rationales are locally sensitive to state changes. This enables efficient and targeted exploration without requiring additional exploration models, thereby improving sample efficiency.
- We introduce a global-attribution-based prior attention mechanism that embeds feature-level global attributions from a pretrained policy into the new decision model, accelerating convergence with negligible additional parameters.
- We validate AGRL on multiple autonomous-driving decision-making tasks and show that it achieves faster convergence and competitive final performance compared with representative baseline methods.
2. Related Work
2.1. Active Exploration in Reinforcement Learning
2.2. Attributing Decisions to Inputs
3. Methodology
3.1. Problem Formulation
3.2. Attribution-Guided Reinforcement Learning Framework
3.2.1. Prior-Attention-Enhanced KAN Decision Model
3.2.2. LRP-Based Real-Time Attribution Method
3.2.3. Attribution-Sensitivity-Guided Exploration
| Algorithm 1 Attribution-Sensitivity-Guided Training in AGRL |
|
4. Implementation
4.1. Scenario Modeling
4.2. Observation and Action Space
4.3. Action Space
4.4. Reward Function
- (1)
- Speed reward. To encourage efficient travel, the speed reward is defined aswhere is the current speed of the ego vehicle, and and are the lower and upper bounds of the desired speed interval.
- (2)
- Comfort cost. To discourage excessive longitudinal acceleration and unnecessary lane changes, the comfort cost is given bywhere denotes the longitudinal acceleration and is a binary indicator for lane-change execution.
- (3)
- Safe-distance cost. To penalize unsafe following behavior and near-collision situations, the safe-distance cost is defined aswhere represents the distance to the nearest leading vehicle in the current lane, and controls the penalty strength.
- (4)
- Collision penalty. To explicitly discourage unsafe behaviors that lead to crashes, the collision penalty is defined as
- (5)
- Merging speed penalty. For the main-road yielding scenario, an additional penalty is imposed when the merging vehicle moves too slowly, encouraging the ego vehicle to create sufficient space for safe merging:where denotes the instantaneous speed of the merging vehicle and is the upper speed bound.
5. Experimental Results
5.1. Baselines
- Interpretable value-based RL baseline: KAN-IDRL is included to distinguish the contribution of the KAN-based interpretable policy representation from that of the proposed attribution-guided exploration mechanism. Similar to the above value-based baselines, KAN-IDRL also relies on random exploration.
- Active-exploration baselines: Random Network Distillation (RND) is included as a representative intrinsic-motivation-based exploration method, Intrinsic Curiosity Module (ICM) [13] is added as a representative curiosity-driven exploration method, while SimHash [40] is employed as a representative count-based exploration method. For a fair comparison, RND, ICM, and SimHash are all re-implemented on the same Dueling DQN backbone, with only the exploration mechanism changed. These three methods are selected as baselines to compare the proposed AGRL framework against classical active-exploration strategies.
5.2. Evaluation Metrics
5.3. Policy Evaluation
5.3.1. Decision Model Training
5.3.2. Decision Model Testing
5.4. Ablation Study
5.4.1. Component Ablation
5.4.2. Ablation of Attribution-Based Exploration Signals
5.5. Interpretability Analysis in a Continuous Semantic Driving Scenario
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Wang, X.; Qi, X.; Wang, P.; Yang, J. Decision making framework for autonomous vehicles driving behavior in complex scenarios via hierarchical state machine. Auton. Intell. Syst. 2021, 1, 10. [Google Scholar] [CrossRef] [Scilit]
- Fu, Y.; Li, C.; Yu, F.R.; Luan, T.H.; Zhang, Y. Hybrid autonomous driving guidance strategy combining deep reinforcement learning and expert system. IEEE Trans. Intell. Transp. Syst. 2021, 23, 11273–11286. [Google Scholar] [CrossRef] [Scilit]
- Bae, S.H.; Joo, S.H.; Pyo, J.W.; Yoon, J.S.; Lee, K.; Kuc, T.Y. Finite state machine based vehicle system for autonomous driving in urban environments. In Proceedings of the 2020 20th International Conference on Control, Automation and Systems (ICCAS); IEEE: New York, NY, USA, 2020; pp. 1181–1186. [Google Scholar]
- Omeiza, D.; Webb, H.; Jirotka, M.; Kunze, L. Explanations in autonomous driving: A survey. IEEE Trans. Intell. Transp. Syst. 2021, 23, 10142–10162. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Bai, Y.; Cai, P.; Wen, L.; Fu, D.; Zhang, B.; Yang, X.; Cai, X.; Ma, T.; Guo, J.; et al. Towards knowledge-driven autonomous driving. arXiv 2023, arXiv:2312.04316. [Google Scholar] [CrossRef] [Scilit]
- Jia, X.; Yang, Z.; Li, Q.; Zhang, Z.; Yan, J. Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving. Adv. Neural Inf. Process. Syst. 2024, 37, 819–844. [Google Scholar]
- Sun, J.; Kim, J. Modelling two-dimensional driving behaviours at unsignalised intersection using multi-agent imitation learning. Transp. Res. Part C Emerg. Technol. 2024, 165, 104702. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z.; Sheng, Z.; Chen, S. PE-RLHF: Reinforcement Learning with Human Feedback and physics knowledge for safe and trustworthy autonomous driving. Transp. Res. Part C Emerg. Technol. 2025, 179, 105262. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Jing, P.; Li, Y.; Wang, X.; Wang, Y. Joint optimization of vehicle platoon and traffic signal with mixed traffic flow at intersections: Deep reinforcement learning approach. Transp. Res. Part C Emerg. Technol. 2025, 177, 105184. [Google Scholar] [CrossRef] [Scilit]
- Zhou, R.; Huang, J.; Li, M.; Li, H.; Cao, H.; Song, X. Knowledge transfer from simple to complex: A safe and efficient reinforcement learning framework for autonomous driving decision-making. Adv. Eng. Inform. 2025, 65, 103188. [Google Scholar] [CrossRef] [Scilit]
- Yang, K.; Tao, J.; Lyu, J.; Li, X. Exploration and anti-exploration with distributional random network distillation. arXiv 2024, arXiv:2401.09750. [Google Scholar] [CrossRef] [Scilit]
- Shyam, P.; Jaśkowski, W.; Gomez, F. Model-based active exploration. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2019; pp. 5779–5788. [Google Scholar]
- Pathak, D.; Agrawal, P.; Efros, A.A.; Darrell, T. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2017; pp. 2778–2787. [Google Scholar]
- Houthooft, R.; Chen, X.; Duan, Y.; Schulman, J.; De Turck, F.; Abbeel, P. Vime: Variational information maximizing exploration. Adv. Neural Inf. Process. Syst. 2016, 29. [Google Scholar]
- Huang, J.; Zhou, R.; Li, M.; Li, H.; Liu, Y.; Song, X. From black-box to white-box: Interpretable deep reinforcement learning with Kolmogorov-Arnold networks for autonomous driving. Transp. Res. Part C Emerg. Technol. 2026, 182, 105386. [Google Scholar] [CrossRef] [Scilit]
- Strehl, A.L.; Littman, M.L. An analysis of model-based interval estimation for Markov decision processes. J. Comput. Syst. Sci. 2008, 74, 1309–1331. [Google Scholar] [CrossRef] [Scilit]
- Azar, M.G.; Osband, I.; Munos, R. Minimax regret bounds for reinforcement learning. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2017; pp. 263–272. [Google Scholar]
- Bellemare, M.; Srinivasan, S.; Ostrovski, G.; Schaul, T.; Saxton, D.; Munos, R. Unifying count-based exploration and intrinsic motivation. Adv. Neural Inf. Process. Syst. 2016, 29, 1479–1487. [Google Scholar]
- Ostrovski, G.; Bellemare, M.G.; Oord, A.; Munos, R. Count-based exploration with neural density models. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2017; pp. 2721–2730. [Google Scholar]
- Machado, M.C.; Bellemare, M.G.; Bowling, M. Count-based exploration with the successor representation. In Proceedings of the AAAI Conference on Artificial Intelligence; IEEE: New York, NY, USA, 2020; Volume 34, pp. 5125–5133. [Google Scholar]
- Lobel, S.; Bagaria, A.; Konidaris, G. Flipping coins to estimate pseudocounts for exploration in reinforcement learning. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2023; pp. 22594–22613. [Google Scholar]
- Burda, Y.; Edwards, H.; Storkey, A.; Klimov, O. Exploration by random network distillation. arXiv 2018, arXiv:1810.12894. [Google Scholar] [CrossRef] [Scilit]
- Lundberg, S.M.; Lee, S.I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30; Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: New York, NY, USA, 2017; pp. 4765–4774. [Google Scholar]
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should i trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; IEEE: New York, NY, USA, 2016; pp. 1135–1144. [Google Scholar]
- Baehrens, D.; Schroeter, T.; Harmeling, S.; Kawanabe, M.; Hansen, K.; Müller, K.R. How to explain individual classification decisions. J. Mach. Learn. Res. 2010, 11, 1803–1831. [Google Scholar]
- Khorram, S.; Lawson, T.; Fuxin, L. iGOS++ integrated gradient optimized saliency by bilateral perturbations. In Proceedings of the Conference on Health, Inference, and Learning; IEEE: New York, NY, USA, 2021; pp. 174–182. [Google Scholar]
- Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic attribution for deep networks. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2017; pp. 3319–3328. [Google Scholar]
- Smilkov, D.; Thorat, N.; Kim, B.; Viégas, F.; Wattenberg, M. Smoothgrad: Removing noise by adding noise. arXiv 2017, arXiv:1706.03825. [Google Scholar] [CrossRef] [Scilit]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 618–626. [Google Scholar]
- Chattopadhay, A.; Sarkar, A.; Howlader, P.; Balasubramanian, V.N. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2018; pp. 839–847. [Google Scholar]
- Montavon, G.; Binder, A.; Lapuschkin, S.; Samek, W.; Müller, K.R. Layer-wise relevance propagation: An overview. In Explainable AI: Interpreting, Explaining and Visualizing Deep Learning; Springer: Berlin/Heidelberg, Germany, 2019; pp. 193–209. [Google Scholar]
- Li, J.; Zhang, C.; Zhou, J.T.; Fu, H.; Xia, S.; Hu, Q. Deep-LIFT: Deep label-specific feature learning for image annotation. IEEE Trans. Cybern. 2021, 52, 7732–7741. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Schaul, T.; Hessel, M.; Hasselt, H.; Lanctot, M.; Freitas, N. Dueling network architectures for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning; PMLR: Tokyo, Japan, 2016; pp. 1995–2003. [Google Scholar]
- Leurent, E. An Environment for Autonomous Driving Decision-Making. 2018. Available online: https://github.com/eleurent/highway-env (accessed on 1 March 2026).
- Bellotti, F.; Lazzaroni, L.; Capello, A.; Cossu, M.; De Gloria, A.; Berta, R. Explaining a deep reinforcement learning (DRL)-based automated driving agent in highway simulations. IEEE Access 2023, 11, 28522–28550. [Google Scholar] [CrossRef] [Scilit]
- Treiber, M.; Hennecke, A.; Helbing, D. Congested traffic states in empirical observations and microscopic simulations. Phys. Rev. E 2000, 62, 1805. [Google Scholar] [CrossRef] [Scilit]
- Kesting, A.; Treiber, M.; Helbing, D. General lane-changing model MOBIL for car-following models. Transp. Res. Rec. 2007, 1999, 86–94. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI Conference on Artificial Intelligence; IEEE: New York, NY, USA, 2016; Volume 30. [Google Scholar]
- Tang, H.; Houthooft, R.; Foote, D.; Stooke, A.; Xi Chen, O.; Duan, Y.; Schulman, J.; DeTurck, F.; Abbeel, P. #Exploration: A study of count-based exploration for deep reinforcement learning. Adv. Neural Inf. Process. Syst. 2017, 30, 2750–2759. [Google Scholar]









| Feature Type | Variable | Standard Deviation |
|---|---|---|
| Binary/discrete | presence | 0 |
| Continuous | x | 0.03 |
| Continuous | y | 0.03 |
| Continuous | 0.02 | |
| Continuous | 0.02 | |
| Continuous | 0.01 |
| Parameters | Highway Lane-Change | Main-Road Yielding | On-Ramp Merging |
|---|---|---|---|
| Total Episodes | 300 | 1000 | 2000 |
| Discount Factor | 0.8 | 0.8 | 0.95 |
| Soft Update Rate | 0.005 | 0.005 | 0.005 |
| Exploration Decay Rate | |||
| Replay Buffer Capacity | |||
| Batch Size | 256 | 256 | 256 |
| Scenario | Window Size W | Reward Threshold |
|---|---|---|
| Highway lane-change | 30 | 24.0 |
| Main-road yielding | 100 | 9.0 |
| On-ramp merging | 200 | 7.5 |
| Method | Training Collisions | Convergence Episodes |
|---|---|---|
| DQN | ||
| DDQN | ||
| Dueling DQN | ||
| KAN-IDRL | ||
| RND | ||
| SimHash | ||
| ICM | ||
| AGRL |
| Method | Trainable Params. | Runtime/Step | Total Training Time | Inference Latency |
|---|---|---|---|---|
| Dueling DQN | 76,292 | 0.0600 | 3099.96 s | 0.0589 |
| KAN-IDRL | 2400 | 0.0754 | 4086.24 s | 0.0625 |
| RND | 184,452 | 0.0616 | 3148.73 s | 0.0587 |
| SimHash | 76,292 | 0.0606 | 3175.52 s | 0.0589 |
| ICM | 251,911 | 0.0620 | 3130.67 s | 0.0588 |
| AGRL | 2436 | 0.0830 | 4759.39 s | 0.0624 |
| Method | Training Collisions | Convergence Episodes |
|---|---|---|
| DQN | ||
| DDQN | ||
| Dueling DQN | ||
| KAN-IDRL | ||
| RND | ||
| SimHash | ||
| ICM | ||
| AGRL |
| Method | Training Collisions | Convergence Episodes |
|---|---|---|
| DQN | ||
| DDQN | ||
| Dueling DQN | ||
| KAN-IDRL | ||
| RND | ||
| SimHash | ||
| ICM | ||
| AGRL |
| Method | Avg. Cumulative Reward | Success Rate (%) | Avg. Speed (m/s) | Safety Cost (×) | Comfort Cost (×) |
|---|---|---|---|---|---|
| DQN | 26.04 ± 1.84 | 96.35 ± 2.54 | 23.89 ± 1.12 | −8.5 ± 8.21 | −6.1 ± 0.81 |
| DDQN | 26.13 ± 1.70 | 95.11 ± 2.89 | 23.84 ± 1.08 | −7.4 ± 7.68 | −5.3 ± 0.57 |
| Dueling DQN | 25.86 ± 1.76 | 95.72 ± 2.67 | 24.30 ± 1.15 | −9.3 ± 8.46 | −4.6 ± 0.74 |
| KAN | 25.49 ± 1.58 | 95.67 ± 2.31 | 23.40 ± 0.96 | −7.8 ± 7.24 | −4.3 ± 0.62 |
| RND | 26.25 ± 1.69 | 95.34 ± 2.76 | 23.76 ± 1.04 | −9.6 ± 8.73 | −5.4 ± 0.79 |
| SimHash | 25.59 ± 1.63 | 95.76 ± 2.48 | 23.84 ± 1.01 | −2.4 ± 2.16 | −6.3 ± 0.86 |
| ICM | 26.18 ± 1.72 | 95.58 ± 2.62 | 23.91 ± 1.06 | −8.8 ± 8.05 | −5.7 ± 0.77 |
| AGRL | 26.57 ± 1.21 | 98.45 ± 2.18 | 23.21 ± 0.88 | −2.3 ± 1.74 | −2.7 ± 0.61 |
| Method | Avg. Cumulative Reward | Success Rate (%) | Avg. Speed (m/s) | Merging Speed Cost (×) | Safety Cost (×) |
|---|---|---|---|---|---|
| DQN | 10.07 ± 0.46 | 96.45 ± 2.18 | 29.95 ± 0.42 | −1.58 ± 0.61 | −0.68 ± 0.29 |
| DDQN | 10.03 ± 0.51 | 95.40 ± 2.47 | 29.92 ± 0.45 | −2.01 ± 0.74 | −0.79 ± 0.33 |
| Dueling DQN | 10.29 ± 0.43 | 95.53 ± 2.36 | 29.81 ± 0.48 | −1.51 ± 0.58 | −0.84 ± 0.36 |
| KAN | 9.649 ± 0.57 | 95.11 ± 2.69 | 29.03 ± 0.63 | −2.29 ± 0.82 | −1.24 ± 0.45 |
| RND | 10.11 ± 0.49 | 96.24 ± 2.21 | 29.94 ± 0.44 | −2.07 ± 0.76 | −0.87 ± 0.37 |
| SimHash | 10.19 ± 0.47 | 95.97 ± 2.33 | 29.75 ± 0.50 | −1.89 ± 0.69 | −0.90 ± 0.39 |
| ICM | 10.16 ± 0.48 | 96.08 ± 2.29 | 29.88 ± 0.46 | −1.96 ± 0.72 | −0.86 ± 0.35 |
| AGRL | 10.32 ± 0.31 | 97.67 ± 1.42 | 29.96 ± 0.28 | −2.09 ± 0.64 | −0.69 ± 0.27 |
| Method | Avg. Cumulative Reward | Success Rate (%) | Avg. Speed (m/s) | Safety Cost (×) | Comfort Cost (×) |
|---|---|---|---|---|---|
| DQN | 8.32 ± 0.54 | 97.13 ± 1.96 | 29.71 ± 0.39 | −2.15 ± 0.82 | −9.62 ± 1.74 |
| DDQN | 8.51 ± 0.49 | 96.35 ± 2.21 | 29.84 ± 0.35 | −1.65 ± 0.71 | −5.84 ± 1.21 |
| Dueling DQN | 8.12 ± 0.61 | 93.51 ± 2.84 | 29.74 ± 0.42 | −1.86 ± 0.76 | −8.59 ± 1.58 |
| KAN | 8.16 ± 0.58 | 96.87 ± 2.05 | 29.65 ± 0.44 | −2.31 ± 0.89 | −6.48 ± 1.36 |
| RND | 8.04 ± 0.63 | 97.05 ± 1.98 | 29.55 ± 0.47 | −2.19 ± 0.85 | −6.74 ± 1.42 |
| SimHash | 8.42 ± 0.52 | 96.75 ± 2.13 | 29.21 ± 0.51 | −1.98 ± 0.79 | −7.51 ± 1.49 |
| ICM | 8.39 ± 0.55 | 96.92 ± 2.07 | 29.68 ± 0.43 | −1.91 ± 0.77 | −6.21 ± 1.30 |
| AGRL | 8.64 ± 0.34 | 98.32 ± 1.17 | 29.57 ± 0.31 | −1.24 ± 0.52 | −4.34 ± 0.96 |
| Model | Mean Euclidean Distance | Mean Cosine Similarity | Mean Sign-Flip Rate | Mean Top-5 Overlap |
|---|---|---|---|---|
| Well-trained | 1.83 | 0.79 | 18.5% | 73.3% |
| Untrained | 1.44 | 0.48 | 35.2% | 53.3% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Huang, J.; Zhou, R.; Wang, Y.; Song, X. Attribution-Guided Active Exploration in Deep Reinforcement Learning for Autonomous Driving Decision-Making. Appl. Sci. 2026, 16, 4931. https://doi.org/10.3390/app16104931
Huang J, Zhou R, Wang Y, Song X. Attribution-Guided Active Exploration in Deep Reinforcement Learning for Autonomous Driving Decision-Making. Applied Sciences. 2026; 16(10):4931. https://doi.org/10.3390/app16104931
Chicago/Turabian StyleHuang, Jiakun, Rongliang Zhou, Yanlong Wang, and Xiaolin Song. 2026. "Attribution-Guided Active Exploration in Deep Reinforcement Learning for Autonomous Driving Decision-Making" Applied Sciences 16, no. 10: 4931. https://doi.org/10.3390/app16104931
APA StyleHuang, J., Zhou, R., Wang, Y., & Song, X. (2026). Attribution-Guided Active Exploration in Deep Reinforcement Learning for Autonomous Driving Decision-Making. Applied Sciences, 16(10), 4931. https://doi.org/10.3390/app16104931

