arXiv · 2609.35426
Frontier Learning: Training LLM Reasoners at the Edge of Capability
Abstract
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi. 2026-09-28. Frontier Learning: Training LLM Reasoners at the Edge of Capability. https://arxiv.org/abs/2609.35426
Cite the original work for its findings. Save a collection to share your selection of sources.