arXiv · 1205.4217
Thompson Sampling: An Asymptotically Optimal Finite Time Analysis
Abstract
The question of the optimality of Thompson Sampling for solving the stochastic multi-armed bandit problem had been open since 1933. In this paper we answer it positively for the case of Bernoulli rewards by providing the first finite-time analysis that matches the asymptotic rate given in the Lai and Robbins lower bound for the cumulative regret. The proof is accompanied by a numerical comparison with other optimal policies, experiments that have been lacking in the literature until now for the Bernoulli case.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Emilie Kaufmann, Nathaniel Korda, Rémi Munos. 2012-07-19. Thompson Sampling: An Asymptotically Optimal Finite Time Analysis. https://arxiv.org/abs/1205.4217
Cite the original work for its findings. Save a collection to share your selection of sources.