arXiv · 2610.04837
CARET: Training-Free Test-Time Scaling for Repository-Level Code Completion
Abstract
Retrieval-augmented generation (RAG) dominates repository-level code completion: it retrieves cross-file context (R), then decodes one greedy completion (G). Existing work mainly focuses on retrieval and stops there. We argue both stages can be improved together, with generation in particular gaining from test-time scaling. We present CARET, a training-free method. For R, CARET routes among retrieval contexts using the agreement among its own samples, cascading to an alternative context when the samples scatter. For G, it samples candidates over a cached prompt prefix, so the long retrieved context is encoded once rather than once per sample. It then selects the final completion by reverse-context likelihood: a correct completion makes the code after the cursor more probable, so the same model grades its own candidates by reading ahead. Across CrossCodeEval, RepoEval-Line, and RepoEval-API with six code models (1.1B to 7B, four families), CARET improves exact match in all 18 combinations by 10.8 points on average over greedy decoding and 5.3 over self-consistency@10. Token-level compute stays near 1.45 times one generation (measured wall-clock 1.3-2.9 times, growing with generator size). Improving retrieval and generation together yields more accurate code than improving retrieval alone, at a budget that stays close to a single pass.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiajie Wang, Yutong Zhao, Kebin Peng, Sen He, Qing Guo, Tianlin Li. 2026-10-04. CARET: Training-Free Test-Time Scaling for Repository-Level Code Completion. https://arxiv.org/abs/2610.04837
Cite the original work for its findings. Save a collection to share your selection of sources.