arXiv · 2606.03077
Libra: Efficient Resource Management for Agentic RL Post-Training
Abstract
Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the rollout stage generates trajectories while invoking tools, producing long-tailed and non-stationary workloads that expose two fundamental challenges. First, due to the long-tailed response distribution, a small fraction of trajectories dominates rollout makespan.Second, rollout and training differ in their compute patterns, memory demands, and sensitivity to sequence length. As the policy evolves, shifts in the workload distribution further change their relative resource demands, making it difficult to maintain balanced execution across the two stages. We present Libra, an adaptive runtime for agentic RL post-training with two complementary components: (1) intra-stage scheduling via a Causality-Guided Bucket Scheduler that routes requests across execution buckets with different parallelism configurations, reducing delays from rollout stragglers; and (2) cross-stage coordination that dynamically reallocates workers between rollout and training as the workload changes. It moves workers between the two stages through a non-blocking protocol without interrupting ongoing training. Evaluated on a 48x NVIDIA A800 GPU cluster and a 160x Ascend 910B3 NPU cluster across three agentic benchmarks, Libra achieves up to 4.2x higher throughput and up to 2.7x faster reward convergence
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kaiwen Chen, Xin Tan, Jingzong Li, Zhi Zhou, Cen Li, Jiang Liu, Meng Jie, Jiazhi Jiang, Hong Xu. 2026-09-16. Libra: Efficient Resource Management for Agentic RL Post-Training. https://arxiv.org/abs/2606.03077
Cite the original work for its findings. Save a collection to share your selection of sources.