Optimal Sample Complexity of Stable Discounted Markov Decision Processes
We study the optimal sample complexity of tabular reinforcement learning for infinite-horizon discounted Markov decision processes. The unrestricted minimax sample complexity is known to scale as $\widetildeΘ((1-γ)^{-3}ε^{-2})$, where $γ$ denotes the discount factor and $ε$ is the solution-error tolerance. However, this worst-case rate does not account for stability structures that commonly arise in operational environments. Using the variance-reduced Q-learning framework, we show that, without imposing any additional stability assumptions, both sup-norm Q-function estimation and policy learning have leading sample-complexity dependence $\widetildeΘ(H(1-γ)^{-2}ε^{-2})$, where $H=|v^*|_{\mathrm{span}}$ is the span of the optimal value function. Moreover, assuming a uniform mixing time upper bound $t_{\mathrm{mix}}$ over all policies, the optimal Q-function can be estimated up to a constant shift with sample complexity $\widetildeΘ\left(t_{\mathrm{mix}}^3ε^{-2}\right)$, independent of $(1-γ)^{-1}$. Matching lower bounds establish the sharpness of these leading dependencies and reveal a separation between estimation and control: although the Q-function can be learned up to an additive constant at a horizon-free rate, policy learning generally retains its $(1-γ)^{-2}$ dependence.