Search arXivSearch

arXiv · 2602.05591

Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets

Abstract

Robust Markov decision processes (MDPs) have attracted significant interest due to their ability to protect MDPs from poor out-of-sample performance in the presence of ambiguity. In contrast to classical MDPs, which account for stochasticity by modeling the dynamics through a stochastic process with a known transition kernel, a robust MDP additionally accounts for ambiguity by optimizing against the most adverse transition kernel from an ambiguity set constructed via historical data. In this paper, we develop a unified solution framework for a broad class of robust MDPs with $s$-rectangular ambiguity sets, where the most adverse transition probabilities are considered independently for each state. Using our algorithms, we show that $s$-rectangular robust MDPs with $1$- and $2$-norm as well as $ϕ$-divergence ambiguity sets can be solved several orders of magnitude faster than with state-of-the-art commercial solvers, and often only a logarithmic factor slower than classical MDPs. We demonstrate the favorable scaling properties of our algorithms on a range of synthetically generated as well as standard benchmark instances.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chin Pang Ho, Marek Petrik, Wolfram Wiesemann. 2026-02-05. Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets. https://arxiv.org/abs/2602.05591

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Control of chaos with minimal information transfer

This paper studies set-invariance and stabilization of hyperbolic sets over rate-limited channels. Our main results reveal a phenomenon which cannot be seen from a linearized analysis: the smallest data rate above which a hyperbolic set $Q$ can be made invariant is bounded below by the difference between two measures of instability: the first one describing the total instability on $Q$, and the second one describing the intrinsic instability which does not lead to exit from $Q$. In rigorous terms, these two quantities are the sum of unstable Lyapunov exponents and the metric entropy of an associated bundle random dynamical system, respectively. The gap between the two is well-known in dynamical systems and is often related to escape rates. A vanishing gap corresponds to the existence of a strange attractor inside $Q$ supporting an SRB measure. In this case, no information transfer to the controller is necessary, because the attractor already guarantees invariance. We prove that our lower bound is tight in two extreme cases, the one just described and the one without intrinsic instability. Furthermore, we apply our techniques to the problem of local uniform stabilization to a hyperbolic set and discuss an example built on the Hénon horseshoe.

math.OC

Tsallis Entropy Regularization for Linear Quadratic Regulator and Kullback-Leibler Control

Shannon entropy regularization is widely adopted in optimal control due to its ability to promote exploration and enhance robustness, e.g., maximum entropy reinforcement learning known as Soft Actor-Critic. The aim of this paper is to show that formulations based on Tsallis entropy, which is a one-parameter extension of Shannon entropy, retain many of the structural and computational advantages of Shannon-entropy-based approaches while offering additional benefits. In particular, we derive a closed-form solution for the linear quadratic regulator and an efficient computational method for the Kullback-Leibler control problem. We also demonstrate its usefulness in balancing between exploration and sparsity of the obtained control law.

math.OC

Computing the nearest scattering passive system

In this paper, we consider linear time-invariant control systems which are bounded real, also known as scattering passive. Our main theoretical contribution is to show the equivalence between such systems and port-Hamiltonian (PH) systems whose factors satisfy certain linear matrix inequalities. Based on this result, we propose a formulation for the problem of finding the nearest bounded real system to a given system, and design an algorithm combining alternating optimization and Nesterov's fast gradient method. This formulation also allows us to check whether a given system is bounded real by solving a semidefinite program, and provide a PH parametrization for it. We illustrate our proposed algorithms on real-world and synthetic data sets.

math.OC