Regret Analysis of Retry-Based Bandits
We provide the first regret analysis of ReMax in stochastic multi-armed bandits. Originally introduced for reinforcement learning, ReMax is motivated by the role of exploration when multiple attempts are allowed, and is closely related to retry-based objectives such as pass@$k$ and max@$k$, which value the best outcome across those attempts. Given a posterior over arm values, ReMax chooses a sampling distribution that maximizes the posterior expected maximum arm value over $M$ virtual draws. For Gaussian rewards with $K$ arms and any fixed integer $M\ge2$, we characterize optimal sampling distributions through an expected-improvement balance condition and prove an $O(\sqrt{(M-1)KT\log T})$ finite-time bound on expected regret. Our analysis separates a logarithmic contribution from rounds without optimal-arm underestimation and a recovery term that dominates the bound. For every fixed instance with pairwise distinct arm means, we further prove an asymptotic logarithmic regret upper bound whose coefficient is at most $M-1$ times the optimal coefficient by Lai and Robbins. When $M=2$, this bound implies the asymptotic optimality of ReMax. Gaussian-bandit experiments show that ReMax performs competitively with Thompson sampling and KL-UCB. They also reveal a trade-off: larger $M$ reduces regret incurred during optimal-arm underestimation, but this reduction does not necessarily translate into lower total regret. Controlled recovery experiments show that separating suboptimal means accelerates recovery, while severe underestimation of the optimal arm can lead to long delays before it is sampled again.