TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, processing visual observations and language instructions through a large language model before predicting robot actions. Although effective, this design incurs substantial computation and memory overhead. In this work, we introduce TurboVLA, a compact VLA architecture built on a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design directly constructs task-conditioned representations while avoiding the overhead of an LLM-centered execution pathway. On LIBERO, TurboVLA achieves 97.6% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GiB inference VRAM on a consumer-grade RTX 4090. Notably, a 0.4B TurboVLA achieves 88.06% success on RoboTwin 2.0, even matching or outperforming substantially larger VLA policies. These results demonstrate that the simple $V + L \to A$ design of TurboVLA can achieve high performance without requiring an LLM-centric execution pathway, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.