arXiv · 2609.25643
Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces
Abstract
Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang. 2026-09-22. Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces. https://arxiv.org/abs/2609.25643
Cite the original work for its findings. Save a collection to share your selection of sources.