One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy
Vision-language-action (VLA) models can use visual prediction to anticipate future states, but dense visual features make the generative sequence grow with the number of camera views, prediction horizon, and encoder resolution. Whether such dense representations are necessary for effective control remains unclear. We introduce OneWM-VLA, which represents each retained camera view with one predictive token per future step. Adaptive Attention Pooling compresses visual features into compact latents, which are jointly generated with robot actions under a conditional flow-matching objective. Future observations provide the latent targets during training and are not required at inference. This design incorporates visual prediction into a pretrained VLA policy while keeping the generative sequence compact. On MetaWorld~MT50, OneWM-VLA improves the average success rate of the $π_0$ backbone from $47.91\%$ to $61.53\%$, reaching $72.01\%$ after 60k training steps. It also achieves $98.1\%$ success on LIBERO and raises Fold Cloth success on a real Piper arm from $20.0\%$ to $60.0\%$ relative to $π_0$. Comparisons on two additional VLA backbones consistently favor one token over three across the evaluated checkpoints. A matched ablation at a longer action horizon further shows that removing the latent loss reduces success from $58.09\%$ to $21.64\%$, supporting the benefit of future supervision for policy learning.