arXiv · 2402.11877
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
Abstract
Reinforcement learning has witnessed significant advancements, particularly with the emergence of model-based approaches. Among these, $Q$-learning has proven to be a powerful algorithm in model-free settings. However, the extension of $Q$-learning to a model-based framework remains relatively unexplored. In this paper, we investigate the sample complexity of $Q$-learning when integrated with a model-based approach. The proposed algorihtms learns both the model and Q-value in an online manner. We demonstrate a near-optimal sample complexity result within a broad range of step sizes.
Explore related subjects
Keep this discovery
Han-Dong Lim, HyeAnn Lee, Donghwan Lee. 2024-02-19. Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ. https://arxiv.org/abs/2402.11877
Cite the original work for its findings. Save a collection to share your selection of sources.