Next Article in Journal
Rain Detection in Solar Insecticidal Lamp IoTs Systems Based on Multivariate Wireless Signal Feature Learning
Next Article in Special Issue
Reinforcement Learning for Enhancing Bitcoin Risk-Aware Trading with Predictive Signals
Previous Article in Journal
Distributed Photovoltaic–Storage Hierarchical Aggregation Method Based on Multi-Source Multi-Scale Data Fusion
Previous Article in Special Issue
MoE-World: A Mixture-of-Experts Architecture for Multi-Task World Models
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Enhancing Adversarial Policy Learning via Value-Based Reward Shaping

1
Rocket Force Engineering University, Xi’an 710025, China
2
Northwestern Polytechnical University, Xi’an 710072, China
3
Southwestern University of Finance and Economics, Chengdu 611130, China
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(2), 463; https://doi.org/10.3390/electronics15020463
Submission received: 21 December 2025 / Revised: 16 January 2026 / Accepted: 17 January 2026 / Published: 21 January 2026

Abstract

In adversarial reinforcement learning, designing dense reward functions is a traditional approach to address the sparsity of adversarial objectives. However, conventional reward design often relies on high-quality domain knowledge and may fail in practice, thereby inducing objective misalignment—a discrepancy between optimizing the designed reward and achieving the true adversarial utility. To reduce this discrepancy, a Value-Based Reward Shaping (VBRS) framework is proposed. VBRS integrates an intrinsic state-value estimate, which is a dynamic predictor of long-term utility, into the immediate reward function. As a result, exploration can be encouraged toward states predicted to be strategically advantageous, potentially avoiding some local optima in practice. Experiments demonstrate that VBRS outperforms a baseline that relies solely on the original reward function. The results confirm that the proposed method enhances adversarial performance and helps bridge the gap between designed reward guidance and the adversarial objective.
Keywords: reinforcement learning; adversarial reinforcement learning; reward shaping reinforcement learning; adversarial reinforcement learning; reward shaping

Share and Cite

MDPI and ACS Style

Hou, B.; Pan, G.; Chen, Y. Enhancing Adversarial Policy Learning via Value-Based Reward Shaping. Electronics 2026, 15, 463. https://doi.org/10.3390/electronics15020463

AMA Style

Hou B, Pan G, Chen Y. Enhancing Adversarial Policy Learning via Value-Based Reward Shaping. Electronics. 2026; 15(2):463. https://doi.org/10.3390/electronics15020463

Chicago/Turabian Style

Hou, Bo, Guangyu Pan, and Yao Chen. 2026. "Enhancing Adversarial Policy Learning via Value-Based Reward Shaping" Electronics 15, no. 2: 463. https://doi.org/10.3390/electronics15020463

APA Style

Hou, B., Pan, G., & Chen, Y. (2026). Enhancing Adversarial Policy Learning via Value-Based Reward Shaping. Electronics, 15(2), 463. https://doi.org/10.3390/electronics15020463

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop