Search arXiv⌕ Search

arXiv · 2609.16370

Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching

Abstract

Wireless edge caching networks typically consist of many independent Base Stations (BSs), each facing its own request rate and content popularity profile. Training a Reinforcement Learning (RL) caching agent from scratch at every BS forces each agent to relearn, through slow trial and error, a decision problem that is structurally identical across the network. Meta-reinforcement learning removes this redundancy by learning a shared initialization that adapts to any BS in a few local updates; however, meta-training itself becomes the bottleneck at scale: the meta-gradient must be estimated from a small subset of BSs at each meta-iteration, and sampling this subset uniformly at random yields a high-variance estimate, an issue existing meta-RL caching frameworks leave unaddressed. This paper proposes a meta-reinforcement learning framework for caching across independent, non-overlapping BSs that directly targets this bottleneck. Each BS runs a local Proximal Policy Optimization (PPO) agent, formulated as a Semi-Markov Decision Process (SMDP) over content popularity, size, lifetime, and importance, while a shared meta-policy is learned via a Model-Agnostic Meta-Learning (MAML)-style loop. To scale meta-training and accelerate convergence, we introduce gradient-based clustering, which groups BSs by local gradient similarity and draws from every cluster, in proportion to its size, at each meta-iteration. We prove, via an Analysis of Variance (ANOVA)-style decomposition of gradient variance, that this strategy yields a strictly lower-variance meta-gradient estimator than uniform random sampling under BS heterogeneity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Farnaz Niknia, Ping Wang. 2026-09-14. Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching. https://arxiv.org/abs/2609.16370

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

xTRUCE: A Provably Safe Arbiter for Multi-xApp Conflict Mitigation in Agentic O-RAN

The open radio access network (O-RAN) is evolving toward agentic operation, where large language model (LLM)-driven xApps/rApps generate control proposals under operator intents. However, such proposals may be conflicting, infeasible, or hallucinated, and no existing system jointly provides proposal-independent safety, priority-aware reconciliation, and traceable feedback. To this end, we propose a provably safe arbiter, namely xTRUCE, in the near-real-time (Near-RT) RAN intelligent controller for mitigating multi-xApp conflicts in gNB control. We first develop a structured xApp proposal interface and a three-layer constraint hierarchy that places physical limits and operator-defined rules above relaxable performance targets, alongside a dual-timescale control action space. A two-stage arbitration mechanism then minimizes target shortfalls in the operator-priority order to finalize safe E2 actions within the Near-RT latency budget, while returning conflict certificates to xApps and the operator for renegotiation. Finally, we implement xTRUCE in a multi-cell O-RAN use case, and evaluate its multi-process prototype through simulations with live API-backed LLM xApps and over-the-air experiments on OpenAirInterface/FlexRIC-based O-RAN stacks. Results show that xTRUCE ensures gNB control safety with $100\%$ protected services despite severe proposal hallucinations, achieves priority-consistent performance satisfaction under overload, efficiently guides LLM intent renegotiation via certificates, and keeps a delay-safe E2 control loop.

cs.NI↗

Packet iSlip

This paper examines input/output buffered crossbar switches under combined packet and cell data. Our switch architecture uses input buffering with Virtual Output Queues to avoid Head of Line Blocking. The switch fabric is a crossbar with no speedup. We use a modified iSlip [McKeown] crossbar scheduler geared towards packet data, called piSlip. Cell ports are largely unmodified from standard iSlip behavior. For packet output ports, we introduce changes to the grant pointer which minimizes output latency caused by packet reassembly. Our model uses several output states, including packet cut-through. From simulation results, we show that piSlip with virtual cut-through offers latency characteristics significantly better than unmodified iSlip with similar packet port interfaces. Simulation further shows that piSlip and iSlip have similar maximum and average buffering requirements.

cs.NI↗

Dynamic Task and Resource Scheduling Towards Space-Air-Ground-Sea Integrated Network

In the context of 6G ubiquitous connectivity, the space-air-ground-sea integrated network (SAGSIN) emerges as a new paradigm for pervasive service provisioning. To support expanding maritime activities in infrastructure-scarce ocean areas, we propose an innovative dynamic task and resource scheduling approach for SAGSIN to deliver computing services for vessels. It integrates broad-coverage satellites, relay-capable high-altitude platform (HAP), energy-sufficient coastal base station (BS), and flexibly deployed uncrewed aerial vehicles (UAVs) to accommodate wide-area, highly mobile, and sustained maritime services. To address the challenge of task scheduling across four layers, a dynamic task offloading algorithm is developed. It steers task flows toward servers with light loads, strong computing capabilities, and high-rate links based on real-time system states to reduce task execution delay, integrating an anticipatory satellite handover strategy to mitigate post-handover congestion and improving satellite resource utilization. Considering the limited endurance of UAVs, we impose residual energy constraints to ensure task backlog handover and safe return. Furthermore, the UAV-BS bandwidth allocation, UAV trajectories, and computing resource allocation are jointly optimized to enhance the connectivity among low-altitude devices and accelerate task completion. Simulation results validate the proposed method's superior adaptability to system resource variations during task execution in complex maritime environments, achieving at least a 23% reduction in average task delay over benchmarks.

cs.NI↗