arXiv · 2510.25811
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
Abstract
We consider a stochastic multi-armed bandit problem with i.i.d. rewards where the expected reward function is multimodal with at most m modes. We propose the first known computationally tractable algorithm for computing the solution to the Graves-Lai optimization problem, which in turn enables the implementation of asymptotically optimal algorithms for this bandit problem. The code for the proposed algorithms is publicly available at https://github.com/wilrev/MultimodalBandits
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
William Réveillard, Richard Combes. 2025-10-29. Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms. https://arxiv.org/abs/2510.25811
Cite the original work for its findings. Save a collection to share your selection of sources.