arXiv · 2610.07696
ESP: Energy-Score Policy for One-Step Multimodal Action Generation
Abstract
Generative action models based on diffusion and flow matching have been increasingly adopted in vision-language-action (VLA) policies for their ability to capture diverse behaviors, including multiple valid action sequences under the same observation and instruction. Their iterative sampling procedures, however, require repeated network evaluations to generate each action chunk, increasing inference latency in closed-loop control. We propose ESP (Energy-Score Policy), a teacher-free approach that maps policy context and noise directly to an action chunk in a single network evaluation. ESP trains the action head with the energy score rather than mean squared error. Whereas squared-error regression targets the conditional mean, the energy score is strictly proper: its expected value is uniquely minimized by the target distribution. This provides a principled objective for learning multimodal action distributions without iterative sampling, with exact recovery at the population optimum when the model can represent the target distribution. Experiments on both simulation and real-world manipulation tasks demonstrate competitive task success with substantially lower action-generation latency than the flow matching baseline. These results support direct distributional learning as an efficient alternative to iterative generative robot policies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lilika Makabe, Heecheol Kim, Yasuyuki Matsushita. 2026-10-06. ESP: Energy-Score Policy for One-Step Multimodal Action Generation. https://arxiv.org/abs/2610.07696
Cite the original work for its findings. Save a collection to share your selection of sources.