Search arXiv⌕ Search

arXiv · 2609.33988

Regime-Dependent Value of CVaR in Preventive Maintenance Scheduling under RUL Uncertainty

Abstract

We study preventive-maintenance scheduling for a small fleet when several assets compete for a limited number of maintenance slots and their remaining useful lives (RULs) are uncertain. The planning horizon is divided into discrete periods; each asset may be maintained at most once, and at most $K$ assets may be maintained in any one period. Future usage and RUL errors are represented by scenarios. We compare a risk-neutral policy that minimizes expected cost with a risk-aware policy that minimizes expected cost plus conditional value-at-risk (CVaR), where $\operatorname{CVaR}_{0.90}$ is the mean cost among the worst 10\% of scenarios. Both policies are solved exactly by enumerating all capacity-feasible schedules. A $3\times5$ experiment combines capacities $K=1,2,3$ with five increasing levels of RUL uncertainty, with ten paired replications for each combination. The risk-aware policy provides its largest benefit when RUL uncertainty is low-to-moderate and more than one maintenance action can be executed per period. In the best observed case, it reduces CVaR by 13.4\% with an expected-cost increase of only 0.2 relative cost units. At high uncertainty, the benefit becomes negligible and the two policies can select identical schedules. The results show that risk-aware scheduling creates operational value only when prognostic information is sufficiently informative and the maintenance system has enough flexibility to act on it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jerzy Baranowski, Waldemar Bauer. 2026-09-27. Regime-Dependent Value of CVaR in Preventive Maintenance Scheduling under RUL Uncertainty. https://arxiv.org/abs/2609.33988

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates

The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetration. We propose the use of reinforcement learning (RL) to optimize the initialization of NR, and introduce a quantum-enhanced RL environment update mechanism that addresses the combinatorially large action space at each RL timestep by formulating the voltage adjustment task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, solved with an Ising machine. RL initialization is benchmarked against flat start and start from the DC (linearized) PF solution on a standard 4-bus system, Iwamoto's ill-conditioned 11-bus system, and the IEEE 118-bus system under normal and stressed loading and reactive power limits, with verified operational solutions. On all systems, a supervised initializer refined by RL requires fewer NR iterations than flat and DC starts and than the same initializer without RL, for all seeds. For example, on the 118-bus system under normal and stressed loading, it reached 2.04 and 2.86 NR iterations, compared with 3.02 and 5.13 from DC start and 2.61 and 3.09 without RL. In wall-clock time, this pays off only for an initializer integrated into the solver and reused for many solves on a fixed topology. On the 4-bus system, a quantum-enhanced RL agent with a quantum-inspired annealer moved challenging initial states that required 29 and 44 NR iterations to initializations that required three NR iterations within one RL timestep.

eess.SY↗

Safety for Weakly-Hard Control Systems via Graph-Based Barrier Functions

Despite significant advancement in technology, communication and computational failures are still prevalent in safety-critical engineering applications. Often, networked control systems experience packet dropouts, leading to open-loop behavior that significantly affects the behavior of the system. Similarly, in real-time control applications, control tasks frequently experience computational overruns and thus occasionally no new actuator command is issued. This article addresses the safety verification and controller synthesis problem for a class of control systems subject to weakly-hard constraints, i.e., a set of window-based constraints where the number of failures are bounded within a given time horizon. The results are based on a new notion of graph-based barrier functions that are specifically tailored to the considered system class, offering a set of constraints whose satisfaction leads to safety guarantees despite such failures. Subsequent reformulations of the safety constraints are proposed to alleviate conservatism and improve computational tractability, and the resulting trade-offs are discussed. Finally, several numerical case studies including linear and polynomial systems demonstrate the effectiveness of the proposed approach.

eess.SY↗

Model-Free Output Feedback Stabilization via Policy Gradient Methods

Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that have been successfully applied to solve problems pertaining to unknown linear dynamical systems, the policy gradient (PG) method stands out due to its ease of implementation and can solve the problem in a model-free manner. However, most of the existing works on PG methods for unknown linear dynamical systems assume full-state feedback. In this paper, we take a step towards model-free learning for partially observed linear dynamical systems with output feedback and focus on the fundamental stabilization problem of the system. We propose an algorithmic framework that stretches the boundary of PG methods to the problem without global convergence guarantees. We show that by leveraging zeroth-order PG update based on system trajectories and its convergence to stationary points, the proposed algorithms return a stabilizing output feedback policy for discrete-time linear dynamical systems. We also explicitly characterize the sample complexity of our algorithm and verify the effectiveness of the algorithm using numerical examples.

eess.SY↗