Skip What You Can Predict: Predictive Repositioning for Policy Optimization for Efficient LLM Training
Reinforcement learning with verifiable rewards (RLVR) can improve the reasoning ability of large language models, but repeatedly updating a policy on the same rollout batch is expensive: every additional update requires another backward pass, and multi-step methods pay for all intermediate optimization steps. We introduce Predictive Repositioning for Policy Optimization (PrePO), which uses two observed optimizer transitions to estimate a farther point along the same-batch optimization trajectory, moves partway toward that point, evaluates the original objective there, and applies a corrective update. This gives PrePO a fixed active cost that does not grow with the virtual optimization depth. Our analysis gives finite-horizon error bounds for AdamW and Muon and shows that sufficiently accurate endpoint estimates preserve the usual descent and convergence behavior of smooth gradient descent. In RLVR experiments, PrePO reaches matched performance targets in fewer optimization steps and lower wall-clock time than the corresponding baselines. We further evaluate the same update mechanism in supervised fine-tuning on a different dataset, showing that its use is not restricted to the original RLVR setting. Together, these results suggest that PrePO provides a practical mechanism for approximating repeated same-batch optimization while avoiding the cost of explicitly executing every intermediate update.