arXiv · 2609.32727
Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication
Abstract
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_δ^{\ast})ε^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_δ^{\ast})ε^{-2})$ for learning an $ε$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_δ^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_δ^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chenyu Lu, Zijun Chen, Nian Si. 2026-09-26. Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication. https://arxiv.org/abs/2609.32727
Cite the original work for its findings. Save a collection to share your selection of sources.