Search arXiv⌕ Search

arXiv · 2609.33637

Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss

Abstract

In this paper, we propose a hierarchical multi-agent reinforcement learning framework for coordinating robot teams in warehouse environments under communication loss. We partition the robot team into groups, with centralized coordination within each group and distributed coordination across groups. Each group uses a recurrent predictor to estimate unavailable interaction information due to communication loss. A higher-level policy then generates a compact coordination reference that conditions the local control policy within each group. A predictive safety filter evaluates and modifies the proposed controls when they violate safety constraints. Simulation results show improved task completion under communication loss, reduced communication growth as the team size increases, and safe operation in the tested scenarios.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Weihao Sun, Gehui Xu, Andreas A. Malikopoulos. 2026-09-27. Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss. https://arxiv.org/abs/2609.33637

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Output-Positive Adaptive Control of Parabolic PDE-ODE Cascades

In this paper, we propose a safe adaptive boundary control strategy for a class of parabolic partial differential equation-ordinary differential equation (PDE-ODE) cascaded systems with parametric uncertainties in both the PDE and ODE subsystems. The proposed design is built upon an adaptive Control Barrier Function (aCBF) framework that incorporates high-relative-degree CBFs together with a batch least-squares identification (BaLSI)-based adaptive control that guarantees exact parameter identification in finite time. The proposed controller ensures the positivity, i.e., the safety, of the plant output state that is the furthest state from the control input, as well as the exponential regulation of the overall plant state to zero. Numerical simulations are provided to demonstrate the effectiveness of the proposed approach.

eess.SY↗

Grid-ECO: Grid Aware Electric Vehicle Charging Stations Placement Optimizer

We develop Grid-ECO, a method for optimally allocating electric vehicle charging stations (EVCS) within a distribution feeder while accounting for EV charging demand at census-level granularity. The underlying problem is a mixed-integer bilinear program (MIBLP) and requires satisfying nonlinear, nonconvex, three-phase unbalanced AC network constraints while including integer siting and sizing decision variables. Existing works cannot guarantee AC feasibility or optimality without either i) relaxing the integer decision variable space or ii) convexifying AC constraints. Grid-ECO solves the exact MIBLP to near-zero optimality gap while prioritizing candidate charging locations using grid voltage and current sensitivity metrics. To solve the MIBLP exactly, we leverage the global optimization algorithm: spatial branch-and-bound (sBnB). To scale the approach to large-scale feeders, we develop a presolving routine that combines i) a penalty-based NLP heuristic for warm-start with b) a sequential bound tightening (SBT) algorithm to derive tight bounds on both bilinear and lifted McCormick variables. Case studies using realistic Seattle city data demonstrate that Grid-ECO substantially outperforms an off-the-shelf commercial sBnB solver. In three of the four test cases, the commercial solver failed to identify a feasible solution or certify optimality within the 12-hour time limit, whereas Grid-ECO solved all instances to a reported optimality gap of 0.00%, with a maximum solver time of 753s. The results further show that tightening both bilinear and lifted McCormick variables reduces sBnB node exploration by up to 98% and solution time by up to 69%, while preserving AC-feasible solutions across all test cases.

eess.SY↗

SCORE: Statistical Certification of Regions of Attraction via Extreme Value Theory

Certifying the Region of Attraction (ROA) for high-dimensional nonlinear dynamical systems remains a severe computational bottleneck. Traditional deterministic verification methods provide hard guarantees but suffer from the curse of dimensionality, typically failing to scale beyond 20 dimensions. To overcome these limitations, we propose SCORE, a statistical certification framework that shifts from seeking deterministic guarantees to bounding the worst-case safety violation with high statistical confidence. By integrating Projected Stochastic Gradient Langevin Dynamics (PSGLD) with Extreme Value Theory (EVT), we frame ROA certification as a constrained extreme-value estimation problem. Under stationary sampling and a regular local-geometry condition around the global maximum, we show that the Lyapunov derivative belongs to the Weibull maximum domain of attraction. Its finite right endpoint enables statistical estimation of the global maximum of the Lyapunov derivative and construction of an upper confidence bound, conditional on the sampling and inference assumptions. Numerical experiments validate that our EVT-based approach achieves certification tightness competitive to exact Sum of Squares programming on a 2D Van der Pol benchmark. Furthermore, we demonstrate strong scalability by successfully applying the statistical verification procedure to a dense, unstructured 500-dimensional ODE system at a nominal confidence level of 99.99\%, effectively bypassing the severe combinatorial constraints that limit existing formal verification pipelines.

eess.SY↗