arXiv · 1910.07479
Conditional Importance Sampling for Off-Policy Learning
Abstract
The principal contribution of this paper is a conceptual framework for off-policy reinforcement learning, based on conditional expectations of importance sampling ratios. This framework yields new perspectives and understanding of existing off-policy algorithms, and reveals a broad space of unexplored algorithms. We theoretically analyse this space, and concretely investigate several algorithms that arise from this framework.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mark Rowland, Anna Harutyunyan, Hado van Hasselt, Diana Borsa, Tom Schaul, Rémi Munos, Will Dabney. 2020-07-30. Conditional Importance Sampling for Off-Policy Learning. https://arxiv.org/abs/1910.07479
Cite the original work for its findings. Save a collection to share your selection of sources.