Next Article in Journal
High-Temperature Hydrogen Sensing Performance of Ni-Doped TiO2 Prepared by Co-Precipitation Method
Previous Article in Journal
A Unified Fourth-Order Tensor-Based Smart Community System
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Policy-Gradient and Actor-Critic Based State Representation Learning for Safe Driving of Autonomous Vehicles

Department of Electrical, Computer and Biomedical Engineering, Ryerson University, Toronto, ON M5B2K3, Canada
*
Author to whom correspondence should be addressed.
Sensors 2020, 20(21), 5991; https://doi.org/10.3390/s20215991
Submission received: 1 September 2020 / Revised: 8 October 2020 / Accepted: 12 October 2020 / Published: 22 October 2020
(This article belongs to the Section Intelligent Sensors)

Abstract

In this paper, we propose an environment perception framework for autonomous driving using state representation learning (SRL). Unlike existing Q-learning based methods for efficient environment perception and object detection, our proposed method takes the learning loss into account under deterministic as well as stochastic policy gradient. Through a combination of variational autoencoder (VAE), deep deterministic policy gradient (DDPG), and soft actor-critic (SAC), we focus on uninterrupted and reasonably safe autonomous driving without steering off the track for a considerable driving distance. Our proposed technique exhibits learning in autonomous vehicles under complex interactions with the environment, without being explicitly trained on driving datasets. To ensure the effectiveness of the scheme over a sustained period of time, we employ a reward-penalty based system where a negative reward is associated with an unfavourable action and a positive reward is awarded for favourable actions. The results obtained through simulations on DonKey simulator show the effectiveness of our proposed method by examining the variations in policy loss, value loss, reward function, and cumulative reward for ‘VAE+DDPG’ and ‘VAE+SAC’ over the learning process.
Keywords: state representation learning; variational auto encoder; deep deterministic policy gradient; soft actor-critic; autonomous driving; Markov decision process state representation learning; variational auto encoder; deep deterministic policy gradient; soft actor-critic; autonomous driving; Markov decision process

Share and Cite

MDPI and ACS Style

Gupta, A.; Khwaja, A.S.; Anpalagan, A.; Guan, L.; Venkatesh, B. Policy-Gradient and Actor-Critic Based State Representation Learning for Safe Driving of Autonomous Vehicles. Sensors 2020, 20, 5991. https://doi.org/10.3390/s20215991

AMA Style

Gupta A, Khwaja AS, Anpalagan A, Guan L, Venkatesh B. Policy-Gradient and Actor-Critic Based State Representation Learning for Safe Driving of Autonomous Vehicles. Sensors. 2020; 20(21):5991. https://doi.org/10.3390/s20215991

Chicago/Turabian Style

Gupta, Abhishek, Ahmed Shaharyar Khwaja, Alagan Anpalagan, Ling Guan, and Bala Venkatesh. 2020. "Policy-Gradient and Actor-Critic Based State Representation Learning for Safe Driving of Autonomous Vehicles" Sensors 20, no. 21: 5991. https://doi.org/10.3390/s20215991

APA Style

Gupta, A., Khwaja, A. S., Anpalagan, A., Guan, L., & Venkatesh, B. (2020). Policy-Gradient and Actor-Critic Based State Representation Learning for Safe Driving of Autonomous Vehicles. Sensors, 20(21), 5991. https://doi.org/10.3390/s20215991

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop