arXiv · 1808.08763
On the convergence of optimistic policy iteration for stochastic shortest path problem
Abstract
In this paper, we prove some convergence results of a special case of optimistic policy iteration algorithm for stochastic shortest path problem. We consider both Monte Carlo and $TD(\lambda)$ methods for the policy evaluation step under the condition that the termination state will eventually be reached almost surely.
Explore related subjects
Keep this discovery
Yuanlong Chen. 2018-08-27. On the convergence of optimistic policy iteration for stochastic shortest path problem. https://arxiv.org/abs/1808.08763
Cite the original work for its findings. Save a collection to share your selection of sources.