arXiv · 2610.05446
Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning
Abstract
Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level policy while retaining accumulated task rewards. Remaining execution time distinguishes subgoal continuation from a new high-level decision. Together with primitive-action conditioning, it enables off-policy Bellman updates based on the stationary environment transition law. HTB combines these one-step updates with multi-step suffix returns and truncated relabeling, reducing dependence on intermediate value estimates. A shared value component supports learning across actions, while nonnegative residuals constrain upward corrections relative to that component. At a fixed mixture weight of 0.95, HTB achieves 32.8% AntFall success versus 9.6% for matched local HIRO over five paired seeds at 10M environment steps. Ablations identify contributions from recursive continuation and mixed supervision; fixed-policy tests show more accurate predictions for actions whose returns were excluded from fitting.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bingyun Liu, Yuheng Jing. 2026-10-04. Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning. https://arxiv.org/abs/2610.05446
Cite the original work for its findings. Save a collection to share your selection of sources.