arXiv · 2511.05187
Gradient Prediction with Control Variates in the Cheap-Forward Regime
Abstract
We study whether otherwise-idle inference resources could reduce the scarce-GPU cost of training. Our analysis uses a simulated compute ledger in which fleet work is billed at a fraction of a scarce-GPU forward; all experiments run on a regular GPU. Our algorithm predicts gradients with a reduced-precision, inference-style reverse-mode program and combines many predictions with a few exact gradients through a control variate, so approximation error becomes variance rather than bias. On a 124M-parameter language model and selected short training windows, the method can lower simulated ledger cost relative to the tested baselines when fleet work is sufficiently cheap. Experiments spanning 10M-774M parameters show both transfers and failures. We do not test inference-only hardware, end-to-end distributed latency, or a full optimizer-by-batch-size baseline sweep.
Explore related subjects
Keep this discovery
Kamil Ciosek, Nicolò Felicioni, Juan Elenter, Ehsan Imani. 2026-09-02. Gradient Prediction with Control Variates in the Cheap-Forward Regime. https://arxiv.org/abs/2511.05187
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.