Next Article in Journal
Kinematic Model and Redundant Space Analysis of 4-DOF Redundant Robot
Next Article in Special Issue
PFVAE: A Planar Flow-Based Variational Auto-Encoder Prediction Model for Time Series Data
Previous Article in Journal
Novel Generalized Proportional Fractional Integral Inequalities on Probabilistic Random Variables and Their Applications
Previous Article in Special Issue
Effect of a Novel Tooth Pitting Model on Mesh Stiffness and Vibration Response of Spur Gears
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Analysis of Performance Measure in Q Learning with UCB Exploration

1
Credit Suisse Securities, New York, NY 10010-3698, USA
2
Zu Chongzhi Center for Mathematics and Computational Sciences, Duke Kunshan University, Kunshan 215316, China
*
Authors to whom correspondence should be addressed.
Most work was done while at Carnegie Mellon University. Opinions expressed in this paper are of the author, and do not reflect the view of Credit Suisse.
Mathematics 2022, 10(4), 575; https://doi.org/10.3390/math10040575
Submission received: 3 January 2022 / Revised: 27 January 2022 / Accepted: 5 February 2022 / Published: 12 February 2022
(This article belongs to the Special Issue Mathematical Method and Application of Machine Learning)

Abstract

Compared to model-based Reinforcement Learning (RL) approaches, model-free RL algorithms, such as Q-learning, require less space and are more expressive, since specifying value functions or policies is more flexible than specifying the model for the environment. This makes model-free algorithms more prevalent in modern deep RL. However, model-based methods can more efficiently extract the information from available data. The Upper Confidence Bound (UCB) bandit can improve the exploration bonuses, and hence increase the data efficiency in the Q-learning framework. The cumulative regret of the Q-learning algorithm with an UCB exploration policy in the episodic Markov Decision Process has recently been explored in the underlying environment of finite state-action space. In this paper, we study the regret bound of the Q-learning algorithm with UCB exploration in the scenario of compact state-action metric space. We present an algorithm that adaptively discretizes the continuous state-action space and iteratively updates Q-values. The algorithm is able to efficiently optimize rewards and minimize cumulative regret.
Keywords: reinforcement learning; Q-learning; multi-armed bandit; theory of machine learning reinforcement learning; Q-learning; multi-armed bandit; theory of machine learning

Share and Cite

MDPI and ACS Style

Ye, W.; Chen, D. Analysis of Performance Measure in Q Learning with UCB Exploration. Mathematics 2022, 10, 575. https://doi.org/10.3390/math10040575

AMA Style

Ye W, Chen D. Analysis of Performance Measure in Q Learning with UCB Exploration. Mathematics. 2022; 10(4):575. https://doi.org/10.3390/math10040575

Chicago/Turabian Style

Ye, Weicheng, and Dangxing Chen. 2022. "Analysis of Performance Measure in Q Learning with UCB Exploration" Mathematics 10, no. 4: 575. https://doi.org/10.3390/math10040575

APA Style

Ye, W., & Chen, D. (2022). Analysis of Performance Measure in Q Learning with UCB Exploration. Mathematics, 10(4), 575. https://doi.org/10.3390/math10040575

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop