arXiv · 2610.00604
MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Abstract
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference $π_{0.5}$ baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 $\pm$ 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev. 2026-09-30. MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation. https://arxiv.org/abs/2610.00604
Cite the original work for its findings. Save a collection to share your selection of sources.