arXiv · 1912.10316
Exploring TD error as a heuristic for $σ$ selection in Q($σ$, $λ$)
Abstract
In the landscape of TD algorithms, the Q($σ$, $λ$) algorithm is an algorithm with the ability to perform a multistep backup in an online manner while also successfully unifying the concepts of sampling with using the expectation across all actions for a state. $σ\in [0, 1]$ indicates the extent to which sampling is used. Selecting the value of σ can be based on characteristics of the current state rather than having a constant value or being time based. This report explores the viability of such a TD-error based scheme.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Abhishek Nan. 2019-12-21. Exploring TD error as a heuristic for $σ$ selection in Q($σ$, $λ$). https://arxiv.org/abs/1912.10316
Cite the original work for its findings. Save a collection to share your selection of sources.