Search arXiv⌕ Search

arXiv · 2610.01344

Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers

Abstract

Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yashdeep Chaudhary, Roberto Armellin, Harry Holt. 2026-10-01. Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers. https://arxiv.org/abs/2610.01344

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Using Less for More: When Warm-Starting Accelerates Branch-and-Cut for Stochastic Programs

Two-stage stochastic programs quickly become intractable as the number of scenarios grows. Motivated by this, we propose TULIP, a modular and easy-to-implement three-step warm-start framework for two-stage stochastic (mixed-)integer programs with an exponential number of cuts separated during branch-and-cut. TULIP (a) builds a cheap surrogate of the full problem by reducing the scenario set or by decoupling the two stages, (b) solves it up to a first incumbent to collect the tight cuts separated along the way, and (c) injects them to warm-start the original problem. In short: we use less (a cheaper surrogate) for more (the original problem). Using this modular setup, we propose four methods within this framework, each with a slightly different setting. Across four case studies, we show that this acceleration is governed by a single mechanism, the root cut loop, and we specify it through a closed-form equation. This TULIP speedup model predicts a speedup when the time saved in the root cut loop exceeds the surrogate overhead. In our experiments, a TULIP variant achieves mean speedups of up to 2.56, with gains increasing with the scenario count. In the remaining case studies, TULIP provides little or no runtime benefit, which the TULIP speedup model mostly explains through insufficient root cut loop savings compared to the surrogate overhead.

math.OC↗

Convergence Analysis of the Wasserstein Proximal Algorithm beyond Geodesic Convexity

The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent progress in studying the training dynamics of the noisy gradient descent algorithm on two-layer neural networks in the mean-field regime, we provide in this paper a simple and self-contained analysis for the convergence of the general-purpose Wasserstein proximal algorithm without assuming geodesic convexity of the objective functional. Under a natural Wasserstein analog of the Euclidean Polyak-Łojasiewicz inequality, we establish that the proximal algorithm achieves an unbiased and linear convergence rate. Our convergence rate improves upon existing rates of the proximal algorithm for solving Wasserstein gradient flows under strong geodesic convexity. We also extend our analysis to the inexact proximal algorithm for geodesically semiconvex objectives. In our numerical experiments, proximal training demonstrates a faster convergence rate than the noisy gradient descent algorithm on mean-field neural networks.

math.OC↗

Technological foundations of management decision-making in the reconstruction of complex gas pipeline system

This monograph presents a comprehensive analysis of the technological foundations of management decision-making in the reconstruction of complex gas pipeline systems. The study addresses the challenges posed by the aging infrastructure of gas supply networks and explores advanced strategies to improve their reliability, efficiency, and automation. Particular attention is given to the reconstruction of pipelines with various configurations linear, looped, and parallel systems under non-stationary gas flow conditions. The proposed models and methodologies offer solutions for optimizing operational parameters, improving emergency valve response, and ensuring uninterrupted gas supply through advanced management systems and data-driven decision support tools. Emphasis is placed on the integration of modern technologies, system theory, and feedback mechanisms in the design and operation of reconstructed pipeline systems. This work is intended for engineers, system designers, and researchers in the fields of gas supply, systems engineering, and energy infrastructure.

math.OC↗