Search arXivSearch

arXiv · 2201.10197

Online Actuator Selection and Controller Design for Linear Quadratic Regulation with Unknown System Model

Abstract

We study the simultaneous actuator selection and controller design problem for linear quadratic regulation with Gaussian noise over a finite horizon of length $T$ and unknown system model. We consider both episodic and non-episodic settings of the problem and propose online algorithms that specify both the sets of actuators to be utilized under a cardinality constraint and the controls corresponding to the sets of selected actuators. In the episodic setting, the interaction with the system breaks into $N$ episodes, each of which restarts from a given initial condition and has length $T$. In the non-episodic setting, the interaction goes on continuously. Our online algorithms leverage a multiarmed bandit algorithm to select the sets of actuators and a certainty equivalence approach to design the corresponding controls. We show that our online algorithms yield $\sqrt{N}$-regret for the episodic setting and $T^{2/3}$-regret for the non-episodic setting. We extend our algorithm design and analysis to show scalability with respect to both the total number of candidate actuators and the cardinality constraint. We numerically validate our theoretical results.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lintao Ye, Ming Chi, Zhi-Wei Liu, Vijay Gupta. 2024-09-13. Online Actuator Selection and Controller Design for Linear Quadratic Regulation with Unknown System Model. https://arxiv.org/abs/2201.10197

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A proximal augmented Lagrangian method for nonconvex optimization with equality and inequality constraints

We propose an inexact proximal augmented Lagrangian method (P-ALM) for nonconvex structured optimization problems. The proposed method features an easily implementable rule not only for updating the penalty parameters, but also for adaptively tuning the proximal term. It allows the penalty parameter to grow rapidly in the early stages to speed up progress, while ameliorating the issue of ill-conditioning in later iterations, a well-known drawback of the traditional approach of linearly increasing the penalty parameters. A key element in our analysis lies in the observation that the augmented Lagrangian can be controlled effectively along the iterates, provided an initial feasible point is available. Our analysis, while simple, provides a new theoretical perspective about P-ALM and, as a by-product, results in similar convergence properties for its non-proximal variant, the classical augmented Lagrangian method (ALM). Numerical experiments, including convex and nonconvex problem instances, demonstrate the effectiveness of our approach.

math.OC

Policy Iteration for Stationary Discounted Hamilton--Jacobi--Bellman Equations: A Viscosity Approach

We study policy iteration (PI) for deterministic infinite-horizon discounted control problems characterized by stationary Hamilton--Jacobi--Bellman equations. For general viscosity solutions, the classical gradient-based policy improvement step need not be defined pointwise. We introduce a semi-discrete formulation with centered difference quotients at scale $h$ and a separate artificial-viscosity term of order $O(h)$. The resulting stencil is monotone, and the positive discount yields a resolvent contraction. Under bounded Lipschitz data and a globally Lipschitz minimizing policy map, we prove monotone and geometric convergence of the value iterates for each fixed $h>0$, together with a local quadratic estimate whose constant is of order $h^{-2}$. Under the additional condition $λ>\Lip_x(f)$, we establish $\|V^h-V\|_\infty\le C\sqrt h$ and combine the discretization and iteration errors into a quantitative bound. A bounded Lipschitz example shows that the $\sqrt h$ exponent is sharp for this scheme. The combined estimate gives a sufficient iteration count of order $h^{-1}\log(1/h)$ to attain an error of order $\sqrt h$. In bounded-domain experiments, the smooth one-dimensional benchmark exhibits the predicted discretization plateau, while a nonlinear two-dimensional manufactured benchmark isolates convergence to the discrete solution. Exact policy evaluation gives substantially faster local convergence than the global geometric bound. A neural evaluation diagnostic illustrates the importance of controlling boundary errors as well as interior residuals.

math.OC

A novel L-shaped refinement chain cuts method for two-stage stochastic programs

This paper introduces the L-shaped refinement chain cuts method, a novel decomposition approach for solving two-stage stochastic programs. The proposed method generalizes the classical L-shaped framework by exploiting refinement chains of scenarios, where sequences of increasingly refined scenario partitions are considered. At each refinement level, the full scenario set is partitioned into subgroups, and one subproblem is solved for each subgroup rather than for each individual scenario as in the classical L-shaped method. The proposed framework provides a unified decomposition scheme in which the classical multi-cut and single-cut L-shaped formulations arise as special cases. Theoretical convergence properties to the optimal solution of the original two-stage stochastic program are established at each refinement level. Furthermore, the relationships between consecutive refinement levels are characterized in terms of Benders cuts, leading to the development of an iterative refinement-based solution algorithm. The effectiveness of the proposed method is evaluated on a two-stage stochastic fixed-charge multicommodity network design problem under a mean-risk formulation. Computational experiments validate the benefits of the proposed approach and highlight its applicability to large-scale risk-averse stochastic optimization problems.

math.OC