Search arXivSearch

arXiv · 2309.17262

Estimation and Inference in Distributional Reinforcement Learning

Abstract

In this paper, we study distributional reinforcement learning from the perspective of statistical efficiency. We investigate distributional policy evaluation, aiming to estimate the complete return distribution (denoted $η^π$) attained by a given policy $π$. We use the certainty-equivalence method to construct our estimator $\hatη^π$, given a generative model is available. In this circumstance we need a dataset of size $\widetilde O\left(\frac{|\mathcal{S}||\mathcal{A}|}{\varepsilon^{2p}(1-γ)^{2p+2}}\right)$ to guarantee the $p$-Wasserstein metric between $\hatη^π$ and $η^π$ less than $\varepsilon$ with high probability. This implies the distributional policy evaluation problem can be solved with sample efficiency. Also, we show that under different mild assumptions a dataset of size $\widetilde O\left(\frac{|\mathcal{S}||\mathcal{A}|}{\varepsilon^{2}(1-γ)^{4}}\right)$ suffices to ensure the Kolmogorov metric and total variation metric between $\hatη^π$ and $η^π$ is below $\varepsilon$ with high probability. Furthermore, we investigate the asymptotic behavior of $\hatη^π$. We demonstrate that the ``empirical process'' $\sqrt{n}(\hatη^π-η^π)$ converges weakly to a Gaussian process in the space of bounded functionals on Lipschitz function class $\ell^\infty(\mathcal{F}_{\text{W}})$, also in the space of bounded functionals on indicator function class $\ell^\infty(\mathcal{F}_{\text{KS}})$ and bounded measurable function class $\ell^\infty(\mathcal{F}_{\text{TV}})$ when some mild conditions hold. Our findings give rise to a unified approach to statistical inference of a wide class of statistical functionals of $η^π$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Liangyu Zhang, Yang Peng, Jiadong Liang, Wenhao Yang, Zhihua Zhang. 2024-09-19. Estimation and Inference in Distributional Reinforcement Learning. https://doi.org/10.1214/25-aos2527

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Combinatorial Inference on the Optimal Assortment in Multinomial Logit Models

Assortment optimization has received active explorations in the past few decades due to its practical importance. Despite the extensive literature dealing with optimization algorithms and latent score estimation, uncertainty quantification for the optimal assortment still needs to be explored and is of great practical significance. Instead of estimating and recovering the complete optimal offer set, decision-makers may only be interested in testing whether a given property holds true for the optimal assortment, such as whether they should include several products of interest in the optimal set, or how many categories of products the optimal set should include. This paper proposes a novel inferential framework for testing such properties. We consider the widely adopted multinomial logit (MNL) model, where we assume that each customer will purchase an item within the offered products with a probability proportional to the underlying preference score associated with the product. We reduce inferring a general optimal assortment property to quantifying the uncertainty associated with the sign change point detection of the marginal revenue gaps. We show the asymptotic normality of the marginal revenue gap estimator, and construct a maximum statistic via the gap estimators to detect the sign change point. By approximating the distribution of the maximum statistic with multiplier bootstrap techniques, we propose a valid testing procedure. We also conduct numerical experiments to assess the performance of our method.

stat.ML

C-Learner: Constrained Learning for Causal Inference

Debiasing methods such as augmented inverse propensity weighting (AIPW), and targeted maximum likelihood estimation (TMLE) enjoy asymptotic properties like semiparametric efficiency and double robustness, but can produce unstable estimates in practice that require ad hoc adjustments (e.g., truncating propensity scores). In contrast, simple plug-ins can remain stable but lack these asymptotic guarantees. To achieve the best of both worlds---a plug-in that enjoys strong asymptotic guarantees---we propose a constrained learning framework that trains a nuisance model to minimize prediction error subject to the constraint that the estimated first-order error of the resulting plug-in is zero. To compare different debiasing methods that share the same classical limit, we study a stylized high-dimensional regression problem where nuisance estimation errors do not vanish asymptotically. Our unified analysis covers both $d n$, as well as ridge regularization, and characterizes how overlap affects the estimators' limiting distributions. Under sufficient overlap, our estimator has smaller asymptotic variance than AIPW and TMLE, whereas when overlap deteriorates so much that AIPW and TMLE are no longer root-$n$ consistent, constrained learning still retains the direct plug-in's root-$n$ limit. Empirically, across a range of experimental settings including those with text-based covariates and language models, we observe our estimator outperforms classical debiasing methods in challenging settings with limited overlap between treatment and control, and performs similarly otherwise.

stat.ML

Small Gradient Norm Regret for Online Convex Optimization

This paper introduces a new problem-dependent regret measure for online convex optimization with smooth losses. The notion, which we call the $G^\star$ regret, depends on the cumulative squared gradient norm evaluated at the decision in hindsight. We show that the $G^\star$ regret strictly refines the existing $L^\star$ (small loss) regret, and that it can be arbitrarily sharper when the losses have vanishing curvature around the hindsight decision. We establish upper and lower bounds on the $G^\star$ regret and extend our results to dynamic regret and bandit settings. As a byproduct, we refine the existing convergence analysis of stochastic optimization algorithms in the interpolation regime. Some experiments validate our theoretical findings.

stat.ML