Search arXivSearch

arXiv · 2607.02533

Centralized PPO-Based DRL for Multi-UAV-BS Positioning and Trajectory Optimization in Disaster Response Networks

Abstract

Unmanned aerial vehicle-mounted base stations (UAV-BSs) constitute a flexible and effective solution for global positioning system (GPS)-free emergency and disaster scenarios, where the rapid deployment of communication infrastructure is critical for maximizing life-saving operations. In this work, we extend a centralized learning framework to a multi-UAV-BS network architecture, in which a single centralized UAV-BS -- as an intelligent agent -- coordinates the three-dimensional positioning and navigation of multiple UAV-BSs, while the remaining UAV-BSs actively serve ground user equipments (UEs) with uncertain positions. We formulate a fairness-aware sum-throughput maximization problem for UAV-BS coordination, which is inherently nonconvex due to the non-linear and interference-coupled throughput expressions. To address this challenge, we cast the problem as a Markov Decision Process (MDP) and solve it using a deep reinforcement learning (DRL) framework based on Proximal Policy Optimization (PPO). The central agent interacts with the environment and learns optimal joint positioning policies that guide the serving UAV-BSs to provide efficient, adaptive, and resilient wireless coverage. The proposed approach exploits spatial configuration and radio signal sensing capabilities to dynamically adapt to heterogeneous UE mobility patterns. Extensive simulations are conducted to evaluate the performance of the proposed method. Numerical results demonstrate that PPO shows competitive performance during both training and evaluation phases. Furthermore, comparative analysis with state-of-the-art RL algorithms, namely Deep Deterministic Policy Gradient (DDPG) and Deep QNetwork (DQN), shows that PPO consistently outperforms these methods in terms of convergence stability, mean reward, and network throughput.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Azim Akhtarshenas, Mario Rico Ibanez, Matteo Bernabe, David Lopez-Perez, Merouane Debbah. 2026-06-16. Centralized PPO-Based DRL for Multi-UAV-BS Positioning and Trajectory Optimization in Disaster Response Networks. https://arxiv.org/abs/2607.02533

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Circulant ADMM-Net for Fast High-resolution DoA Estimation

This paper introduces CADMM-Net and CHADMM-Net, two deep neural networks for direction of arrival estimation within the least-absolute shrinkage and selection operator (LASSO) framework. These two networks are based on a structured deep unfolding of the alternating direction method of multipliers (ADMM) algorithm through the use of circulant as well as Hermitian-circulant matrices. Along with a computational complexity of $\mathcal{O}(N\log(N))$ per layer for the inference, where $N$ is the length of the dictionary $\mathbf{A}$, they additionally exhibit a memory footprint of $N$ and approximately half of $N$ for CADMMNet and CHADMM-Net, respectively, compared with $N^{2}$ for ADMM-Net. Furthermore, these structured networks exhibit a competitive performance against ADMM-Net, LISTA, TLISTA, and THLISTA with respect to the detection rate, the angular root-mean square error, and the normalized mean squared error.

eess.SP

BASIIS: Bistatic Angular Sampling and Interpolation for ISAC Setups

Integrated Sensing and Communications (ISAC) is a defining feature of 6G, extending cellular networks with radar-like sensing at limited additional overhead. In bistatic deployments, sensing requires coordinating the transmitter (TX) and receiver (RX) arrays to scan the Cartesian product of angle of departure and arrival, resulting in a four-dimensional sampling problem in the angular domain. This work establishes a complete angular sampling framework for bistatic ISAC, extending the DFT-based optimal-sampling methodology to the full azimuth and elevation domains of both arrays. We show that the bistatic geometry couples the TX and RX elevation angles, and represent this coupling through the ortho-baseline coarray, a virtual array that captures the joint elevation aperture of the array pair. From the coarray we derive a minimal sampling and interpolation scheme, near-lossless and realizable with any beamforming architecture. Monte Carlo simulations confirm the proposed minimal acquisition essentially equalizes the detection accuracy of dense oversampled imaging while acquiring 3 to 5 times fewer TX-RX direction pairs. This allows having bistatic operations with drastically reduced overhead on the radio resource usage of ISAC systems.

eess.SP

Centroid Angle Estimation of Multiple Scatterers Using Monopulse Radar with Frequency Diversity

The monopulse technique determines the angle of a target by comparing signals from two narrow beams, yielding a precise angular estimate with low complexity. However, it struggles to resolve multiple closely spaced scatterers within the same resolution cell. Existing methods for estimating multiple scatterer angles involve complex signal processing and system modifications. We propose an effective method to estimate the angular centroid of scatterers using the mode of monopulse angle estimates. A semi-analytic expression for the angle estimate distribution is derived, confirming that its mode aligns with the centroid. To enhance estimation accuracy, we employ frequency diversity to reduce sample correlation. Numerical results validate the advantages of the proposed method, demonstrating superior performance over conventional techniques with low complexity.

eess.SP