arXiv · 2610.01687
Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
Abstract
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza. 2026-10-01. Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models. https://arxiv.org/abs/2610.01687
Cite the original work for its findings. Save a collection to share your selection of sources.