Search arXivSearch

arXiv · 2111.12316

A note on stabilizing reinforcement learning

Abstract

Reinforcement learning is a general methodology of adaptive optimal control that has attracted much attention in various fields ranging from video game industry to robot manipulators. Despite its remarkable performance demonstrations, plain reinforcement learning controllers do not guarantee stability which compromises their applicability in industry. To provide such guarantees, measures have to be taken. This gives rise to what could generally be called stabilizing reinforcement learning. Concrete approaches range from employment of human overseers to filter out unsafe actions to formally verified shields and fusion with classical stabilizing controllers. A line of attack that utilizes elements of adaptive control has become fairly popular in the recent years. In this note, we critically address such an approach in a fairly general actor-critic setup for nonlinear time-continuous environments. The actor network utilizes a so-called robustifying term that is supposed to compensate for the neural network errors. The corresponding stability analysis is based on the value function itself. We indicate a problem in such a stability analysis and provide a counterexample to the overall control scheme. Implications for such a line of attack in stabilizing reinforcement learning are discussed. Furthermore, unfortunately the said problem possess no fix without a substantial reconsideration of the whole approach. As a positive message, we derive a stochastic critic neural network weight convergence analysis provided that the environment was stabilized.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pavel Osinenko, Grigory Yaremenko, Ilya Osokin. 2022-06-11. A note on stabilizing reinforcement learning. https://arxiv.org/abs/2111.12316

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Effective equidistribution of orbits under semisimple groups on congruence quotients

We prove an effective equidistribution result for periodic orbits of semisimple groups on congruence quotients of an ambient semisimple group.This extends a previous work of Einsiedler, Margulis and Venkatesh. The main new feature is that we allow for periodic orbits of semisimple groups with nontrivial centralizer in the ambient group. Our proof uses crucially an effective closing lemma from work of the author with Lindenstrauss, Margulis,Mohammadi, and Shah.

math.DS

Generalized entropy of measure-induced maps

A classical result by E. Glasner and B. Weiss states that the topological entropy of a map $f$ is zero if and only if the topological entropy of its measure-induced map $f_*$ is zero, where $f_*$ is defined as the push-forward of a measure. In this work, we use generalized entropy to distinguish the complexity of these maps and prove that the measure-induced map is much more complex than the original map. Moreover, we introduce the generalized mean dimension, an invariant that is useful for distinguishing dynamical systems with zero mean dimension, including those with the small-boundary property, and we show a relationship between this new invariant and generalized entropy.

math.DS

The endpoint problem for $\varepsilon$-hypercyclicity

For a fixed $0<\varepsilon<1$, F. Bayart asked in 2024 whether there exists an operator $T$ such that, for every $0<δ<1$, $T$ is $δ$-hypercyclic if and only if $δ\in[\varepsilon,1)$. We answer this question affirmatively by constructing a weighted backward shift on $\ell_2(\mathbb N_0,\ell_2(\mathbb N_0))$ with this property.

math.DS