arXiv · 1709.03570
A KL-LUCB Bandit Algorithm for Large-Scale Crowdsourcing
Abstract
This paper focuses on best-arm identification in multi-armed bandits with bounded rewards. We develop an algorithm that is a fusion of lil-UCB and KL-LUCB, offering the best qualities of the two algorithms in one method. This is achieved by proving a novel anytime confidence bound for the mean of bounded distributions, which is the analogue of the LIL-type bounds recently developed for sub-Gaussian distributions. We corroborate our theoretical results with numerical experiments based on the New Yorker Cartoon Caption Contest.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bob Mankoff, Robert Nowak, Ervin Tanczos. 2017-09-11. A KL-LUCB Bandit Algorithm for Large-Scale Crowdsourcing. https://arxiv.org/abs/1709.03570
Cite the original work for its findings. Save a collection to share your selection of sources.