arXiv · 2607.29078
DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
Abstract
While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead to contexts unfamiliar to the teacher. To improve trajectory quality, recent work on agentic OPD introduces teacher intervention into training rollouts by switching the executor between the student and the teacher. However, existing methods determine how much teacher intervention is needed---but not when. To address this limitation, we propose DASH-OPD (Discrepancy-Aware Switching with Hysteresis for OPD), the first agentic OPD method to perform adaptive, bidirectional executor switching. At each turn, DASH-OPD measures teacher--student discrepancy using a mean log-probability ratio over action tokens. Student-to-teacher ratios on student turns serve as drift signals, while teacher-to-student ratios on teacher turns serve as recovery signals. These signals are accumulated over multiple turns to form drift and recovery evidence, respectively. DASH-OPD switches executors when either type of evidence exceeds its corresponding switching threshold, introducing hysteresis that prevents rapid switching triggered by transient discrepancy fluctuations. Across three benchmarks and two student model sizes, DASH-OPD outperforms five baselines in all 14 task performance comparisons, while requiring the fewest interaction turns in nine of ten efficiency comparisons. Code, models, and training logs are available at https://github.com/Lucian1115/DASH-OPD
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuchen Xia, Qianguo Sun, Chao Song, Junlong Wu, Yiyan Qi, Yunjian Xu. 2026-09-11. DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation. https://arxiv.org/abs/2607.29078
Cite the original work for its findings. Save a collection to share your selection of sources.