Search arXivSearch

arXiv · 2512.08766

Optimal navigation in two-dimensional flows: Control theory and reinforcement learning

Abstract

Zermelo's navigation problem seeks the trajectory of minimal travel time between two points in a fluid flow. We address this problem for an agent -- such as a floating drone or active particle -- that is advected by a two-dimensional flow, self-propels at a fixed speed smaller than or comparable to the characteristic flow velocity, and can steer its direction. The flows considered span increasing levels of complexity, from steady solid-body rotation and time-dependent sink-vortex to the Taylor-Green flow and turbulence in the inverse energy cascade regime. Although optimal-control theory provides time-minimizing trajectories, these solutions become unstable in chaotic regimes characterized by positive finite-time Lyapunov exponents. To design robust navigation strategies, we apply reinforcement learning and compare Q-learning with a one-step actor-critic algorithm. Both methods achieve successful navigation, yielding mean travel times within 3-10% of optimal-control solutions in regular flows, while the discrepancy increases to 35-75% in time-dependent turbulent flows. Finally, we show that agents trained on coarse-grained turbulent flows generalize to the full velocity field. This robustness to incomplete flow information is essential for practical navigation in real-world oceanic and atmospheric environments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vladimir Parfenyev. 2026-06-22. Optimal navigation in two-dimensional flows: Control theory and reinforcement learning. https://doi.org/10.1103/8gbr-vp2r

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Mathematical modeling on peristaltic flow of a Prandtl fluid with effects of slip conditions and inclined magnetic field

The manuscript provides a description of a theoretical analysis of a non-Newtonian Prandtl fluid subject to peristaltic flow through an inclined asymmetric channel. We explore the effect of an inclined magnetic field on the peristaltic flow. This is relevant for applications involving fluid flow in narrow, inclined (tilted) tubes similar to blood vessels or the digestive system. The model also includes thermodynamic aspects such as heat diffusion (the Soret effect) and viscous dissipation resulting from wall-fluid slip conditions, which may help optimize medical devices such as lab-on-a-chip systems and dialysis machines. In this study, the concentration of a generic chemical, temperature, and fluid velocity are taken into account through mass, heat, and momentum balances, respectively. The solution is approximated using numerical techniques suitable for long wavelengths (low frequency) and low Reynolds numbers. The study also discusses trapping phenomena, which are crucial from a clinical point of view. The developed insights can improve the understanding of physiological flows in the gastrointestinal tract and blood vessels. By understanding how the fluid moves and how particles are trapped, these insights may contribute to the design of improved medical pumps and artificial organs. Graphical visualizations are provided for the fluid velocity profile, temperature distribution, and concentration of a generic chemical. Furthermore, the numerical results are validated through comparison with a closed-form solution from a benchmark problem.

physics.flu-dyn

Discovery of a dispersion model at high Peclet numbers

Peclet number characterises the transition from classical Taylor-Aris dispersion to convection-dominated longitudinal solute transport, with the classical model becoming inadequate at extremely high radial Peclet number $Pe_r$. We develop a novel explicit-closure one-dimensional (1-D) effective dispersion model for this high-$Pe_r$ regime by introducing two closure coefficients, $θ_u$ and $θ_d$, whose functional structures are identified using low-frequency transfer-function matching and a modified Kolmogorov-Arnold network (KAN). The resulting model captures the transition from classical Taylor-Aris dispersion at low $Pe_r$ to convection-dominated dispersion at high $Pe_r$. Analysis reveals that, in the high-$Pe_r$ regime, axial transport is redistributed between the effective convection flux and the dispersive flux, resulting in a reduced macroscopic convection velocity. Numerical validation demonstrates close agreement with the convection-diffusion model over the investigated high-$Pe_r$ conditions, while the classical Taylor-Aris model exhibits substantial deviations. Application of the proposed model to averaged flow velocity inversion further demonstrates improved velocity estimation, particularly in the high-$Pe_r$ regime. These results highlight the importance of accounting for non-classical dispersion for reliable contrast-agent-based arterial blood flow velocimetry and provide new insight into high-$Pe_r$ mass transport.

physics.flu-dyn

Optimization of fluid mixing by reinforcement learning using limit cycles of a dynamical system

We propose a method to overcome the difficulties encountered when applying reinforcement learning to fluid mixing processes. The proposed method has two main features: (i) it does not require detailed measurements of the flow state, and (ii) by effectively exploiting a stable limit cycle of a two-dimensional dynamical system (the Li'enard system), it can stably perform optimization without imposing explicit constraints on the control parameters. As an illustrative example, we optimize a process in which a fluid contained in a cylindrical vessel is mixed by periodically rotating the vessel. The resulting optimal vessel motion is physically reasonable: it reverses its direction of rotation before a solid-body rotation state is established. Furthermore, even when the fluid viscosity increases with time during the mixing process, the method can continuously adapt the control parameters to the changing viscosity.

physics.flu-dyn