Search arXivSearch

arXiv · 0905.0463

On ergodic two-armed bandits

Abstract

A device has two arms with unknown deterministic payoffs and the aim is to asymptotically identify the best one without spending too much time on the other. The Narendra algorithm offers a stochastic procedure to this end. We show under weak ergodic assumptions on these deterministic payoffs that the procedure eventually chooses the best arm (i.e., with greatest Cesaro limit) with probability one for appropriate step sequences of the algorithm. In the case of i.i.d. payoffs, this implies a "quenched" version of the "annealed" result of Lamberton, Pagès and Tarrès [Ann. Appl. Probab. 14 (2004) 1424--1454] by the law of iterated logarithm, thus generalizing it. More precisely, if $(η_{\ell,i})_{i\in \mathbb {N}}\in\{0,1\}^{\mathbb {N}}$, $\ell\in\{A,B\}$, are the deterministic reward sequences we would get if we played at time $i$, we obtain infallibility with the same assumption on nonincreasing step sequences on the payoffs as in Lamberton, Pagès and Tarrès [Ann. Appl. Probab. 14 (2004) 1424--1454], replacing the i.i.d. assumption by the hypothesis that the empirical averages $\sum_{i=1}^nη_{A,i}/n$ and $\sum_{i=1}^nη_{B,i}/n$ converge, as $n$ tends to infinity, respectively, to $θ_A$ and $θ_B$, with rate at least $1/(\log n)^{1+\varepsilon}$, for some $\varepsilon >0$. We also show a fallibility result, that is, convergence with positive probability to the choice of the wrong arm, which implies the corresponding result of Lamberton, Pagès and Tarrès [Ann. Appl. Probab. 14 (2004) 1424--1454] in the i.i.d. case.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pierre Tarrès, Pierre Vandekerkhove. 2012-04-26. On ergodic two-armed bandits. https://doi.org/10.1214/10-aap751

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Generalized Edgeworth expansions for integer-valued additive functionals of uniformly elliptic Markov chains

We obtain asymptotic expansions for probabilities $\bbP(S_N=k)$ of partial sums of uniformly bounded integer-valued functionals $\DS S_N=\sum_{n=1}^N f_n(X_n)$ of uniformly elliptic inhomogeneous Markov chains. The expansions involve products of polynomials and trigonometric polynomials, and they hold without additional assumptions. As an application of the explicit formulas of the trigonometric polynomials, we relate existence of the standard Edgeworth expansions of order $r$ to the rate of equidistributions of $S_N$ modulo $m$ for small positive integers $m.$

math.PR

Permutations from Random Walk

Xavier and Yushi run a "random race" as follows. An atomless probability distribution $μ$ on the real line is chosen. The runners begin at zero. At time $i$ Xavier draws $\mathbf{X}_i$ from $μ$ and advances that distance, while Yushi advances by an independent drawing $\mathbf{Y}_i$. After $n$ such moves, what is the probability that Yushi led all the way? That the answer (namely, $4^{-n}\binom{2n}{n}$) is independent of $μ$ follows from a classical theorem of Darling, stating that for symmetric atomless increments, the distribution of each individual rank in the permutation obtained by ranking the partial sums is independent of the step law. We give a self-contained proof and extend the result to the permutations generated by partial sums of uniformly random signed permutations of any fixed, finite, generic set of reals. For atomless increments with mean zero and finite variance, without assuming symmetry, we show that random-walk permutations approach a random object that we call the "Wiener permuton," whose expected pattern densities equal the probabilities of the corresponding permutations generated by finite random walks with centered Laplace increments. Finally, we exhibit an infinite family of constructions whose limiting permutons interpolate between the Wiener permuton and the recursive separable permuton; each has the same intensity permuton, providing a single two-dimensional extension of the classical arcsine law for all of them.

math.PR

On the uniqueness of quasi-stationary distributions for population models with spatial structure

Subcritical population processes are attracted to extinction and do not have non-trivial stationary distributions, which prompts the study of quasi-stationary distributions (QSDs) instead. In contrast to what generally happens for stationary distributions, QSDs may not be unique, even under irreducibility conditions. The general conditions for uniqueness of QSDs are not always easy to check. For the branching process, besides the quasi-limiting distribution there are many other QSDs. In this paper, we investigate whether adding little extra information to the continuous-time branching process is enough to obtain uniqueness. We consider the branching process with genealogy and branching random walks, and show that they have a unique QSD.

math.PR