arXiv · 2610.01023
Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
Abstract
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei. 2026-10-01. Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents. https://arxiv.org/abs/2610.01023
Cite the original work for its findings. Save a collection to share your selection of sources.