Next Article in Journal
Deep Learning-Based Joint CSI Feedback and Hybrid Precoding in FDD mmWave Massive MIMO Systems
Previous Article in Journal
Fast Phylogeny of SARS-CoV-2 by Compression
Previous Article in Special Issue
An Improved Approach towards Multi-Agent Pursuit–Evasion Game Decision-Making Using Deep Reinforcement Learning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment

1
School of Information and Electronics, Hunan City University, Yiyang 413000, China
2
School of Computer Science and Engineering, Central South University, Changsha 410075, China
3
School of Computer Science, National University of Defense Technology, Changsha 410073, China
4
5G&6G Innovation Centre, Department of Electrical and Electronic Engineering, Institute for Communication Systems, University of Surrey, Guildford GU2 7XH, UK
*
Author to whom correspondence should be addressed.
Entropy 2022, 24(4), 440; https://doi.org/10.3390/e24040440
Submission received: 18 January 2022 / Revised: 8 March 2022 / Accepted: 10 March 2022 / Published: 22 March 2022

Abstract

In the field of reinforcement learning, we propose a Correct Proximal Policy Optimization (CPPO) algorithm based on the modified penalty factor β and relative entropy in order to solve the robustness and stationarity of traditional algorithms. Firstly, In the process of reinforcement learning, this paper establishes a strategy evaluation mechanism through the policy distribution function. Secondly, the state space function is quantified by introducing entropy, whereby the approximation policy is used to approximate the real policy distribution, and the kernel function estimation and calculation of relative entropy is used to fit the reward function based on complex problem. Finally, through the comparative analysis on the classic test cases, we demonstrated that our proposed algorithm is effective, has a faster convergence speed and better performance than the traditional PPO algorithm, and the measure of the relative entropy can show the differences. In addition, it can more efficiently use the information of complex environment to learn policies. At the same time, not only can our paper explain the rationality of the policy distribution theory, the proposed framework can also balance between iteration steps, computational complexity and convergence speed, and we also introduced an effective measure of performance using the relative entropy concept.
Keywords: correct proximal policy optimization; approximation theory; reinforcement learning; optimization; policy gradient; entropy correct proximal policy optimization; approximation theory; reinforcement learning; optimization; policy gradient; entropy

Share and Cite

MDPI and ACS Style

Chen, W.; Wong, K.K.L.; Long, S.; Sun, Z. Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment. Entropy 2022, 24, 440. https://doi.org/10.3390/e24040440

AMA Style

Chen W, Wong KKL, Long S, Sun Z. Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment. Entropy. 2022; 24(4):440. https://doi.org/10.3390/e24040440

Chicago/Turabian Style

Chen, Weimin, Kelvin Kian Loong Wong, Sifan Long, and Zhili Sun. 2022. "Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment" Entropy 24, no. 4: 440. https://doi.org/10.3390/e24040440

APA Style

Chen, W., Wong, K. K. L., Long, S., & Sun, Z. (2022). Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment. Entropy, 24(4), 440. https://doi.org/10.3390/e24040440

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop