arXiv · 2609.36371
LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents
Abstract
Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6--81.0% and total verification tokens, which include test generation, by 49.1--62.1% at K=16, while hybrid Best@16 matches or improves on each agent's reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuning Han, Yangchenchen Jin, Tyler Jandreau, Jingwei Sun. 2026-09-28. LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents. https://arxiv.org/abs/2609.36371
Cite the original work for its findings. Save a collection to share your selection of sources.