Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, many existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. We introduce Efficient-WAM, a World-Action Model motivated by the idea that a compact future branch can still support effective action generation, even at lower visual fidelity. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of prioritizing visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. Our 1B-parameter model achieves 98 ms per action chunk during physical deployment, over 30x faster than the Motus baseline with comparable task success.