You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
Generative robot policies trained on demonstrations using behavior cloning often learn actions that are sub-optimal or misaligned with respect to the downstream task. Policy improvement approaches aim to bridge this gap and improve the cumulative reward with minimal interventions on the pre-trained policy. However, an impediment to their practical deployment is the requirement of cumbersome hyperparameter tuning specific to model or task families. In this work, we present an episodic, derivative-free latent policy improvement approach that works with little to no changes across model scales from MLPs to VLAs. We demonstrate that the performance of a pretrained, frozen diffusion or flow matching policy can be improved with respect to a downstream reward by swapping the sampling of initial noise from the prior distribution (typically isotropic Gaussian) with a well-chosen, constant initial noise input---a golden ticket. We show the prevalence of golden tickets by improving the policy performance of $46$ out of $51$ tasks across manipulation benchmarks, with absolute improvements in success rate by up to $79\%$ for simulated tasks, and $28\%$ within $60$ search episodes for real-world tasks. Further, we find that the versatility of our approach opens up new avenues such as simultaneous policy improvement for multiple downstream rewards. Project webpage: https://lottery-tickets.rai-inst.com/.