arXiv · 2208.02409
Randomized Optimal Stopping Problem in Continuous time and Reinforcement Learning Algorithm
Abstract
In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a transformation reduces the optimal stopping problem to a standard optimal control problem. We derive the related HJB equation and prove its solvability. Furthermore, we give a convergence rate of policy iteration and the comparison to classical optimal stopping problem. Based on the theoretical analysis, a reinforcement learning algorithm is designed and numerical results are demonstrated for several models.
Explore related subjects
Keep this discovery
Yuchao Dong. 2022-08-04. Randomized Optimal Stopping Problem in Continuous time and Reinforcement Learning Algorithm. https://arxiv.org/abs/2208.02409
Cite the original work for its findings. Save a collection to share your selection of sources.