Search arXivSearch

arXiv · 2008.11592

Robust Reinforcement Learning: A Case Study in Linear Quadratic Regulation

Abstract

This paper studies the robustness of reinforcement learning algorithms to errors in the learning process. Specifically, we revisit the benchmark problem of discrete-time linear quadratic regulation (LQR) and study the long-standing open question: Under what conditions is the policy iteration method robustly stable from a dynamical systems perspective? Using advanced stability results in control theory, it is shown that policy iteration for LQR is inherently robust to small errors in the learning process and enjoys small-disturbance input-to-state stability: whenever the error in each iteration is bounded and small, the solutions of the policy iteration algorithm are also bounded, and, moreover, enter and stay in a small neighbourhood of the optimal LQR solution. As an application, a novel off-policy optimistic least-squares policy iteration for the LQR problem is proposed, when the system dynamics are subjected to additive stochastic disturbances. The proposed new results in robust reinforcement learning are validated by a numerical example.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bo Pang, Zhong-Ping Jiang. 2021-03-15. Robust Reinforcement Learning: A Case Study in Linear Quadratic Regulation. https://arxiv.org/abs/2008.11592

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Control of chaos with minimal information transfer

This paper studies set-invariance and stabilization of hyperbolic sets over rate-limited channels. Our main results reveal a phenomenon which cannot be seen from a linearized analysis: the smallest data rate above which a hyperbolic set $Q$ can be made invariant is bounded below by the difference between two measures of instability: the first one describing the total instability on $Q$, and the second one describing the intrinsic instability which does not lead to exit from $Q$. In rigorous terms, these two quantities are the sum of unstable Lyapunov exponents and the metric entropy of an associated bundle random dynamical system, respectively. The gap between the two is well-known in dynamical systems and is often related to escape rates. A vanishing gap corresponds to the existence of a strange attractor inside $Q$ supporting an SRB measure. In this case, no information transfer to the controller is necessary, because the attractor already guarantees invariance. We prove that our lower bound is tight in two extreme cases, the one just described and the one without intrinsic instability. Furthermore, we apply our techniques to the problem of local uniform stabilization to a hyperbolic set and discuss an example built on the Hénon horseshoe.

math.OC

Tsallis Entropy Regularization for Linear Quadratic Regulator and Kullback-Leibler Control

Shannon entropy regularization is widely adopted in optimal control due to its ability to promote exploration and enhance robustness, e.g., maximum entropy reinforcement learning known as Soft Actor-Critic. The aim of this paper is to show that formulations based on Tsallis entropy, which is a one-parameter extension of Shannon entropy, retain many of the structural and computational advantages of Shannon-entropy-based approaches while offering additional benefits. In particular, we derive a closed-form solution for the linear quadratic regulator and an efficient computational method for the Kullback-Leibler control problem. We also demonstrate its usefulness in balancing between exploration and sparsity of the obtained control law.

math.OC

Computing the nearest scattering passive system

In this paper, we consider linear time-invariant control systems which are bounded real, also known as scattering passive. Our main theoretical contribution is to show the equivalence between such systems and port-Hamiltonian (PH) systems whose factors satisfy certain linear matrix inequalities. Based on this result, we propose a formulation for the problem of finding the nearest bounded real system to a given system, and design an algorithm combining alternating optimization and Nesterov's fast gradient method. This formulation also allows us to check whether a given system is bounded real by solving a semidefinite program, and provide a PH parametrization for it. We illustrate our proposed algorithms on real-world and synthetic data sets.

math.OC