arXiv · 2601.18952
Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach
Abstract
We propose Kernel Embedding Distributional Reinforcement Learning (KE-DRL) an offline method to estimate kernel mean embeddings of conditional multivariate return distributions from observed offline trajectories with continuous state-action inputs and vector-valued rewards. KE-DRL approximates the conditional embedding and estimates its a finite-dictionary representation coefficients through a maximum mean discrepancy Bellman criterion. For a regular class of return distributions, we show that a Matérn embedding and a Bellman-invariant Sobolev-moment class identifies the unique distributional Bellman fixed point and and derive a Hölder-type bound that controls Wasserstein error by the population embedding Bellman residual. For directly observed responses, we provide finite-sample and uniform error bounds for a regularized conditional mean embedding estimator. Simulation experiments evaluate pointwise embedding recovery against Monte Carlo benchmarks across multiple behavior-target policy pairs. An application to Expedia hotel-search data illustrates conditional multivariate return evaluation under two policies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mehrdad Mohammadi, Qi Zheng, Ruoqing Zhu. 2026-09-19. Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach. https://arxiv.org/abs/2601.18952
Cite the original work for its findings. Save a collection to share your selection of sources.