Search arXivSearch

arXiv · 1907.08591

Zermelo's problem: Optimal point-to-point navigation in 2D turbulent flows using Reinforcement Learning

Abstract

To find the path that minimizes the time to navigate between two given points in a fluid flow is known as Zermelo's problem. Here, we investigate it by using a Reinforcement Learning (RL) approach for the case of a vessel which has a slip velocity with fixed intensity, Vs , but variable direction and navigating in a 2D turbulent sea. We show that an Actor-Critic RL algorithm is able to find quasi-optimal solutions for both time-independent and chaotically evolving flow configurations. For the frozen case, we also compared the results with strategies obtained analytically from continuous Optimal Navigation (ON) protocols. We show that for our application, ON solutions are unstable for the typical duration of the navigation process, and are therefore not useful in practice. On the other hand, RL solutions are much more robust with respect to small changes in the initial conditions and to external noise, even when V s is much smaller than the maximum flow velocity. Furthermore, we show how the RL approach is able to take advantage of the flow properties in order to reach the target, especially when the steering speed is small.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Luca Biferale, Fabio Bonaccorso, Michele Buzzicotti, Patricio Clark Di Leoni, Kristian Gustavsson. 2019-09-26. Zermelo's problem: Optimal point-to-point navigation in 2D turbulent flows using Reinforcement Learning. https://doi.org/10.1063/1.5120370

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Final state sensitivity and fractal basin boundaries from coupled Chialvo neurons

We investigate and quantify the basin geometry and extreme final state uncertainty of two identical electrically asymmetrically coupled Chialvo neurons. The system's diverse behaviors are presented, along with the mathematical reasoning behind its chaotic and nonchaotic dynamics as determined by the structure of the coupled equations. The system is found to be multistable with two qualitatively different attractors. Although each neuron is individually nonchaotic, the chaotic basin takes up the vast majority of the coupled system's state space, but the nonchaotic basin stretches to infinity due to chance synchronization. The boundary between the basins is found to be fractal, leading to extreme final state sensitivity. This uncertainty and its potential effect on the synchronization of biological neurons may have implications for understanding neuronal biology.

nlin.CD

Jordan-Block Degeneracy and Cubic-Order Bifurcating Periodic Orbits in Minimum-Energy Optimal Control of Hamiltonian Equilibria

Equilibria of the Hamiltonian system associated with Pontryagin's minimum principle exhibit an exact doubling of the natural spectrum and, under a simple pairing condition, a Jordan block at every simple purely imaginary eigenvalue. Consequently, the classical Lyapunov Center Theorem does not apply to the augmented system, and no periodic orbit with nonzero optimal control bifurcates at linear order. We establish this mechanism in general and show that an optimal-control-induced periodic family emerges at cubic order in a Lindstedt--Poincaré expansion. The mechanism is illustrated in closed form for the pendulum and evaluated numerically for the planar $L_2$ equilibrium of Hill's restricted three-body problem, where the third-order approximation is validated against an independently computed family of periodic orbits.

nlin.CD

Hypersensitivity and Turnpikes in Optimal Control of Inverted Pendulum: A Dynamical Systems Perspective

The hypersensitivity and turnpike phenomena in the optimal control of an inverted pendulum are investigated from a dynamical-systems perspective. We show that, for a fixed terminal time and a fixed terminal state optimal control problem, (1) the hypersensitivity originates from the fractal structure of the set of initial adjoint variables in the associated Hamiltonian dynamics, (2) the turnpike arises from slow dynamics in the vicinity of a degenerate center manifold, and (3) the escape channels are formed by normally hyperbolic invariant manifolds (NHIMs). As a consequence, small perturbations in the initial adjoint variables lead to qualitatively distinct extremal trajectories, resulting in severe numerical instability in trajectory optimization. Both the fractal structure and the invariant sets are characterized numerically and analytically.

nlin.CD