SearcharxivSearch

arXiv · 2111.12316

A note on stabilizing reinforcement learning

Abstract

Reinforcement learning is a general methodology of adaptive optimal control that has attracted much attention in various fields ranging from video game industry to robot manipulators. Despite its remarkable performance demonstrations, plain reinforcement learning controllers do not guarantee stability which compromises their applicability in industry. To provide such guarantees, measures have to be taken. This gives rise to what could generally be called stabilizing reinforcement learning. Concrete approaches range from employment of human overseers to filter out unsafe actions to formally verified shields and fusion with classical stabilizing controllers. A line of attack that utilizes elements of adaptive control has become fairly popular in the recent years. In this note, we critically address such an approach in a fairly general actor-critic setup for nonlinear time-continuous environments. The actor network utilizes a so-called robustifying term that is supposed to compensate for the neural network errors. The corresponding stability analysis is based on the value function itself. We indicate a problem in such a stability analysis and provide a counterexample to the overall control scheme. Implications for such a line of attack in stabilizing reinforcement learning are discussed. Furthermore, unfortunately the said problem possess no fix without a substantial reconsideration of the whole approach. As a positive message, we derive a stochastic critic neural network weight convergence analysis provided that the environment was stabilized.

Explore related subjects

Keep this discovery

BibTeXRIS

Pavel Osinenko, Grigory Yaremenko, Ilya Osokin. 2021-11-24. A note on stabilizing reinforcement learning. https://arxiv.org/abs/2111.12316

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Admissible Fourier Lengths, KAM Reducibility, and Spectral Applications

We develop a perturbative KAM reducibility theory for one-frequency $\mathrm{SL}(2,\mathbb{R})$ cocycles based on an admissible Fourier length $\ell$. The regularity relevant to the iteration is measured by positive adapted Fourier width rather than ordinary smoothness in the Euclidean length $|n|$. The same length governs Fourier decay, truncation and resonance scales, and the arithmetic condition controlling the small divisors. This framework contains the classical analytic and Gevrey settings, while non-monotone choices of $\ell$ allow classical nowhere differentiable Weierstrass-type perturbations and continuous perturbations outside every positive H\"older class. As spectral applications, we obtain purely absolutely continuous spectrum for every phase and $1/2$-H\"older continuity of the integrated density of states for the associated quasiperiodic Schr\"odinger operators. The Aubry dual has pure point spectrum for Lebesgue almost every dual phase, with eigenfunctions exponentially localized in the metric induced by $\ell$. We also construct nowhere differentiable quasiperiodic potentials with purely absolutely continuous Cantor spectrum.

math.DS

Dynamics inside the attracting basins of some skew products

Polynomial skew products in $\mathbb{C}^2$ are maps of the form $F(z,w)=(P(z),Q(z,w))$, where $P$ and $Q$ are polynomials. Their local dynamics have been widely investigated. In this paper, we study the global dynamics inside Fatou components of some skew products. We consider all the inverse images in a Fatou component of a given point and use the Kobayashi metric to measure the distance between points. In the cases we consider, there are always arbitrarily large Kobayashi balls in the complement of these inverse sets.

math.DS

Ergodicity of dynamical systems without uniqueness of orbits

Recently, there has been considerable interest in the study of non-deterministic dynamical systems. To analyze the chaotic behavior of such systems from a measure-theoretic viewpoint, it is desirable to consider ergodicity. However, the classical definition of ergodicity involves invariant sets, whose definition is not unique for non-deterministic dynamical systems. Thus, we are led to the question of which invariance yields an interesting definition of ergodicity. Here, we propose a definition based on the strong backward invariance and show that analogs of classical results hold. We also consider implications of the Birkhoff ergodic theorem for systems without uniqueness of orbits.

math.DS