arXiv · 2610.05185
Recurrent Latent Visual Search for GUI Grounding
Abstract
GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motivating multi-step visual search. However, existing approaches commonly rely on textual reasoning misaligned with visual space or costly multi-round interactions with external visual tools. To make multi-step visual search an explicit spatial process within the model, we propose ReLaViS, which performs Recurrent Latent Visual Search in a single interaction round. At each step, a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that explicitly represents the search focus. This distribution then aggregates the visual tokens into latent visual evidence, which is recurrently fed back as the next input embedding to condition subsequent search. We further introduce a GUI-aware coarse-to-fine inductive bias through trajectories constructed from flat element annotations, supervising search from the global interface through intermediate element groups to the target. Built on Qwen2.5-VL-7B, ReLaViS improves ScreenSpot-Pro accuracy by 3.1 percentage points to 56.3% with only a 3.5% increase in inference FLOPs and outperforms the matched single-step baseline on all five benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kaiyu Wu, Beichen Zheng, Weiyao Huang, Keze Wang. 2026-10-04. Recurrent Latent Visual Search for GUI Grounding. https://arxiv.org/abs/2610.05185
Cite the original work for its findings. Save a collection to share your selection of sources.