arXiv · 2609.33186
Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations
Abstract
Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Amitakshar Biswas, Yuhan Li, Ruoqing Zhu. 2026-09-27. Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations. https://arxiv.org/abs/2609.33186
Cite the original work for its findings. Save a collection to share your selection of sources.