arXiv · 2610.03372
SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents
Abstract
Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo. 2026-10-02. SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents. https://arxiv.org/abs/2610.03372
Cite the original work for its findings. Save a collection to share your selection of sources.